How to Generate a Sitemap by Scraping a Website
Crawl a site you control, filter discovered URLs to canonical pages, write a valid sitemap, and publish it so search engines can find it.
To generate a sitemap by scraping a website, crawl reachable internal pages within a defined host scope, normalize and deduplicate discovered URLs, keep only preferred canonical pages intended for search discovery, serialize them as XML or a plain URL list, validate the result, and publish it. First check whether your CMS or site software already generates one; Google recommends using the site’s software when available. A sitemap helps search engines discover URLs, but does not guarantee crawling or indexing. Google’s sitemap creation guidance and sitemap overview explain the limits and formats.
1. Check for an existing sitemap
Before building a crawler, check your CMS documentation, common sitemap paths such as /sitemap.xml, and the site’s robots.txt for a Sitemap: line. Your CMS or application may be the more accurate source because it knows which pages are publishable. Crawling is useful when you need an inventory, when the site software does not expose a sitemap, or when you need a separate discovery list for a site you administer.
Do not treat the sitemap as a substitute for access controls or robots rules. It lists discovery candidates; it does not grant permission to crawl them.
2. Define scope and crawl the site
Choose a seed URL and define what counts as internal. For a typical site, restrict the crawler to the same hostname and scheme policy, and decide explicitly whether subdomains belong. Set sensible request concurrency and delays, identify your crawler appropriately, respect the site’s access rules, and avoid crawling URLs that trigger state changes or expensive actions.
For a desktop workflow, Screaming Frog documents this sequence: enter the site URL, run the crawl, then choose Sitemaps > XML Sitemap after it finishes. Its documentation says the free Lite edition supports up to 500 URLs; verify current product limits before relying on them. Its defaults are useful starting points, not universal SEO policy. Screaming Frog’s XML sitemap tutorial.
For a coded workflow, Scrapy spiders begin from URLs, parse responses, and yield further requests; allowed-domain constraints can help set the crawl boundary. Scrapy spider documentation.
3. Normalize, deduplicate, and choose URLs
A crawl produces candidates, not a sitemap policy. Keep one preferred URL for equivalent content. Follow the site’s canonical tags where appropriate and check that the canonical destination is in scope, returns successfully, and is intended as a search landing page.
- Normalize scheme and host consistently, such as HTTPS and the preferred www or non-www form.
- Remove fragments because they identify positions within a document, not separate pages.
- Remove tracking parameters and session identifiers when they do not change the content. Preserve parameters that genuinely identify distinct, indexable pages.
- Exclude redirects, errors, blocked pages, noindex pages, duplicate or canonicalized variants, and pages with no search value according to your site’s policy.
- Check pagination, faceted navigation, PDFs, and alternate-language pages deliberately. Their inclusion depends on the content and sitemap format.
Google advises choosing the canonical URL when equivalent content is available at multiple URLs. Screaming Frog’s default sitemap output, for example, includes internal HTML pages with 200 responses and excludes redirects, errors, robots.txt-blocked pages, noindex pages, canonicalized URLs, paginated URLs, and PDFs. Treat that as an example of tool behavior, not a rule that fits every site. Confirm status, canonical, indexability, duplication, and business intent. Google build guidance.
4. Build an XML sitemap with Python and Scrapy
The following is a compact Scrapy project example. It crawls same-host links from a seed page, avoids fragments and common tracking parameters, records successful HTML pages, and writes an XML sitemap. It is a starting point for a site you control: review canonical tags and robots/noindex policy before publishing. Scrapy honors robots.txt only when configured to do so, so the example enables that setting.
Install and run
python -m pip install scrapy
scrapy runspider sitemap_spider.py -a start_url=https://example.com/ -O sitemap.json
Save this as sitemap_spider.py:
import json
from urllib.parse import urldefrag, urlparse, urlunparse, parse_qsl, urlencode
import scrapy
TRACKING_KEYS = {"gclid", "fbclid", "mc_cid", "mc_eid"}
TRACKING_PREFIXES = ("utm_",)
def normalize_url(url):
url, _fragment = urldefrag(url)
parts = urlparse(url)
if parts.scheme not in ("http", "https"):
return None
# Drop common tracking parameters, preserving other query parameters.
query = [(k, v) for k, v in parse_qsl(parts.query, keep_blank_values=True)
if k.lower() not in TRACKING_KEYS
and not any(k.lower().startswith(p) for p in TRACKING_PREFIXES)]
return urlunparse((parts.scheme.lower(), parts.netloc.lower(),
parts.path or "/", parts.params,
urlencode(query, doseq=True), ""))
class SitemapSpider(scrapy.Spider):
name = "sitemap"
custom_settings = {
"ROBOTSTXT_OBEY": True,
"CONCURRENT_REQUESTS_PER_DOMAIN": 2,
"DOWNLOAD_DELAY": 0.5,
"USER_AGENT": "SitemapInventoryBot/1.0 (+https://example.com/contact)",
"FEEDS": {},
}
def __init__(self, start_url, *args, **kwargs):
super().__init__(*args, **kwargs)
self.start_url = normalize_url(start_url)
if not self.start_url:
raise ValueError("start_url must be an http(s) URL")
self.allowed_domains = [urlparse(self.start_url).hostname]
self.pages = {}
def start_requests(self):
yield scrapy.Request(self.start_url, callback=self.parse)
def parse(self, response):
content_type = response.headers.get(b"Content-Type", b"").decode("latin1").lower()
if response.status == 200 and "text/html" in content_type:
# Default candidate policy: same-host, successful HTML pages.
# Add canonical and robots/noindex checks for your site's policy.
page_url = normalize_url(response.url)
if page_url:
self.pages[page_url] = {"url": page_url, "status": response.status}
for href in response.css("a::attr(href)").getall():
target = normalize_url(response.urljoin(href))
if target and urlparse(target).hostname == urlparse(self.start_url).hostname:
yield scrapy.Request(target, callback=self.parse)
def closed(self, reason):
# Scrapy deduplicates requests by default. Save candidates for review.
with open("sitemap-candidates.json", "w", encoding="utf-8") as f:
json.dump(list(self.pages.values()), f, ensure_ascii=False, indent=2)
This spider intentionally writes candidate records to sitemap-candidates.json; the feed command is optional and is not needed for the crawl. Review the candidates, remove pages that should not be indexed, and apply canonical targets before generating the final sitemap. A production implementation should parse each page’s <link rel="canonical"> and robots directives, verify the canonical target, handle redirects consistently, and record any reliable modification timestamp you already maintain.
Serialize reviewed URLs as XML
After editing the candidate list into a file named urls.txt with one approved absolute URL per line, use this standard-library script to escape XML values and write a sitemap. It checks that every URL is HTTP(S), has a hostname, and shares the selected host. Set EXPECTED_HOST to your canonical hostname.
from urllib.parse import urlparse
from xml.etree.ElementTree import Element, SubElement, ElementTree
EXPECTED_HOST = "www.example.com"
INPUT = "urls.txt"
OUTPUT = "sitemap.xml"
urlset = Element("urlset", xmlns="http://www.sitemaps.org/schemas/sitemap/0.9")
seen = set()
with open(INPUT, encoding="utf-8") as f:
for raw in f:
loc = raw.strip()
if not loc or loc in seen:
continue
parsed = urlparse(loc)
if parsed.scheme not in ("http", "https") or not parsed.hostname:
raise ValueError(f"Not an absolute HTTP(S) URL: {loc}")
if parsed.hostname.lower() != EXPECTED_HOST.lower():
raise ValueError(f"URL outside expected host: {loc}")
seen.add(loc)
SubElement(SubElement(urlset, "url"), "loc").text = loc
ElementTree(urlset).write(OUTPUT, encoding="utf-8", xml_declaration=True)
print(f"Wrote {len(seen)} URLs to {OUTPUT}")
The XML library escapes special characters in text values. For a plain text sitemap, write the same reviewed absolute URLs one per line to a UTF-8 text file. Google supports this simpler format when the sitemap only needs to list page URLs. XML is more versatile when you need image, video, news, or alternate-language extensions. Google’s format guidance.
5. XML fields and sitemap limits
A basic XML sitemap uses the Sitemap protocol’s urlset, url, and loc elements. Add lastmod only when you can provide a consistently accurate modification date. Google says it ignores priority and changefreq; do not add them expecting ranking or crawl benefits, and do not fabricate dates.
A sitemap file is limited to 50,000 URLs or 50 MB uncompressed; larger sites need multiple sitemap files and a sitemap index. Keep all listed URLs absolute and use the preferred canonical form. For specialized content, use the relevant XML extension and follow its requirements.
6. Validate, publish, and submit
- Parse the XML and confirm it has the expected namespace and well-formed entries.
- Check every
locis an absolute URL on the intended host, appears once, and resolves to the preferred canonical page. - Check response status, robots access, noindex directives, and canonical targets against your inclusion policy.
- Publish the file at a stable public URL such as
https://www.example.com/sitemap.xml. - Add
Sitemap: https://www.example.com/sitemap.xmlto the appropriaterobots.txt, or submit the sitemap in Google Search Console. - Review the Search Console Sitemaps report for fetch and processing errors, then keep the sitemap updated as pages change.
A robots.txt file’s rules apply only to the protocol, host, and port where that file is served. A sitemap directive identifies the sitemap location; it does not allow crawling through a disallow rule. See Google’s robots.txt guide and submission guidance.
7. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| URLs appear twice with different hosts or schemes | HTTP/HTTPS or www/non-www variants were not normalized | Choose the site’s canonical host and scheme, then deduplicate after normalization. |
| Tracking or session URLs fill the file | Parameters were treated as distinct pages | Drop known tracking/session parameters; retain parameters only when they represent distinct canonical content. |
| Search Console reports a fetch or parsing error | Bad XML, inaccessible file, invalid URL, or server response issue | Check public access, HTTP status, XML syntax, encoding, and each absolute loc. |
| Disallowed URLs appear in the sitemap | The crawler did not obey robots.txt or the filter policy missed a restriction | Enable robots compliance where appropriate and validate every candidate against the site’s rules. |
| Canonicalized, noindex, or redirected URLs are listed | The crawl output was published without SEO filtering | Resolve canonicals, check indexability and response status, and keep only preferred URLs intended for discovery. |
| Important pages are absent | They are orphaned, require interaction, are not linked from the seed’s reachable pages, or the crawl scope is too narrow | Provide additional seed URLs or use the CMS/database as source of truth; crawling links alone cannot discover unlinked pages. |
| Repeated URLs cause a crawl loop | Query variants, calendar links, or unbounded parameters create an effectively infinite URL space | Normalize parameters, define path/query allowlists, and cap depth or request count. |
| Listed URL returns a soft error page with HTTP 200 | Status alone was used as the inclusion test | Review page content and canonical/indexability signals; a 200 response is not proof the URL belongs in the sitemap. |
8. Performance, reliability, and cost
A sitemap crawl can generate many requests. Limit concurrency and add delay according to the site’s capacity, avoid duplicate fetches, and impose a maximum crawl duration or URL count. Cache or persist crawl state for large inventories so a restart does not begin from zero. For a site you control, an application database or CMS export is often more reliable than discovering all URLs from links because orphaned pages cannot be reached by link traversal.
Keep the generation repeatable: store the scope and exclusion rules in code or configuration, log skipped URLs and reasons, and validate the generated file before replacing the published copy. Deploy atomically so crawlers do not fetch a partially written sitemap. Do not claim a crawl speed or cost without measuring against your own host and setup; request volume, rendering needs, and rate limits vary.
Or skip the browser setup
If your inventory workflow needs rendered page captures for review, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns an image or PDF; it does not crawl a site or generate a sitemap. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers say the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Every feature is available on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Can a sitemap include every URL a crawler finds?
It can, but that is usually not the right selection policy. Include preferred canonical URLs that you want search engines to discover, not every duplicate or parameter variant.
Will submitting a sitemap make pages rank or get indexed?
No. It helps search engines discover URLs, but does not guarantee crawling or indexing.
Should I use XML or a text sitemap?
Use a text file if you only need a list of URLs. Use XML when you need sitemap extensions or structured metadata.
Can crawling find pages that are not linked anywhere?
No. A link-following crawler needs a reachable link or another seed. Use the CMS, database, or known URL sources to find orphaned pages.


