How to Track Competitor Websites at Scale With Sitemap Extraction
Find competitor sitemap URLs, follow sitemap indexes, and compare dated snapshots. Learn where sitemap monitoring helps—and how to verify changes with a crawl.
Sitemap extraction gives you a repeatable way to discover the URLs a competitor publishes and spot changes between collection runs. Find sitemap locations in each site’s robots.txt, recursively fetch sitemap indexes and child files, save each listed URL with its source and collection time, then compare snapshots. Treat a sitemap as a discovery signal: it does not prove a page is live, crawlable, or indexed. Google’s [sitemap guidance](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) explicitly says sitemap submission does not guarantee crawling or indexing.
For useful monitoring, pair sitemap snapshots with a separate, conservative crawl of important URLs. Check response status, redirects, canonical URLs, robots directives, and page content. Sitemap-only monitoring sees only the URLs exposed by the sitemap files you found.
1. Find each competitor’s sitemap
How do I find a competitor’s sitemap?
Start with the canonical host you intend to monitor. Fetch /robots.txt over HTTPS and parse every line beginning with Sitemap:. A robots file may list multiple sitemap locations, including an index. Google’s [robots and sitemap documentation](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) describes sitemap declarations in robots.txt.
curl -fsSL https://www.example.com/robots.txt
If no sitemap directive appears, try common paths such as /sitemap.xml, /sitemap_index.xml, and /sitemap-index.xml. These are discovery fallbacks, not guaranteed protocol paths. Record how you found each file so you can distinguish advertised locations from guesses. If the site redirects between www and non-www, HTTP and HTTPS, or another hostname, record the redirect and decide whether those hosts belong in scope.
Track subdomains separately when they represent separate products, locales, or content collections. A competitor list should include the exact starting host, the final redirected host, the robots URL, and the date you added it.
2. Extract sitemap indexes and child files
How can I extract all URLs from a sitemap index?
A sitemap index contains sitemap locations; a URL-set sitemap contains page locations. Parse the XML root element rather than guessing from a filename. Follow every sitemap location recursively, preserve each source file, and keep going when one child fails so a partial outage does not erase the rest of the run.
Google documents a limit of 50 MB uncompressed or 50,000 URLs per sitemap; larger collections are split across files and organized with an index. The [Sitemaps protocol](https://www.sitemaps.org/protocol.html) also defines sitemap indexes and optional modification metadata. XML should be UTF-8, and special characters in XML values must be escaped.
Runnable Python collector
This script discovers sitemap directives, tries common fallbacks when needed, handles gzip-compressed sitemap responses and .gz files, follows nested indexes, and writes one JSON Lines record per page URL. It also saves a run report containing fetched-file status and errors. Install the dependency with python -m pip install requests, save the script as collect_sitemaps.py, then run python collect_sitemaps.py https://www.example.com.
import gzip
import json
import sys
import time
import xml.etree.ElementTree as ET
from collections import deque
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
import requests
UA = "SitemapResearch/1.0 (+contact your-team@example.com)"
TIMEOUT = (10, 45)
MAX_FILES = 10000
def local_name(tag):
return tag.rsplit("}", 1)[-1].lower()
def child_text(parent, name):
for child in parent:
if local_name(child.tag) == name:
return (child.text or "").strip() or None
return None
def get(session, url):
response = session.get(url, timeout=TIMEOUT, headers={"User-Agent": UA})
response.raise_for_status()
body = response.content
if url.lower().split("?", 1)[0].endswith(".gz"):
body = gzip.decompress(body)
return response, body
def robots_sitemaps(session, robots_url):
response, body = get(session, robots_url)
found = []
for line in body.decode(response.encoding or "utf-8", errors="replace").splitlines():
key, separator, value = line.partition(":")
if separator and key.strip().lower() == "sitemap":
candidate = value.strip()
if candidate:
found.append(urljoin(robots_url, candidate))
return found
def main(start):
parsed = urlparse(start if "://" in start else "https://" + start)
origin = f"{parsed.scheme}://{parsed.netloc}"
session = requests.Session()
started = datetime.now(timezone.utc).isoformat()
errors = []
files = []
page_rows = []
try:
queue = deque(robots_sitemaps(session, origin + "/robots.txt"))
except requests.RequestException as exc:
errors.append({"url": origin + "/robots.txt", "error": str(exc)})
queue = deque()
# Fallback paths are guesses. Keep this list explicit and label them in the report.
if not queue:
for path in ("/sitemap.xml", "/sitemap_index.xml", "/sitemap-index.xml"):
queue.append(origin + path)
seen_files = set()
seen_pages = set()
while queue:
sitemap_url = queue.popleft()
if sitemap_url in seen_files:
continue
if len(seen_files) >= MAX_FILES:
errors.append({"url": sitemap_url, "error": "MAX_FILES limit reached"})
break
seen_files.add(sitemap_url)
try:
response, body = get(session, sitemap_url)
root = ET.fromstring(body)
kind = local_name(root.tag)
files.append({"source": sitemap_url, "status": response.status_code, "kind": kind})
if kind == "sitemapindex":
for item in root:
if local_name(item.tag) == "sitemap":
loc = child_text(item, "loc")
if loc:
queue.append(urljoin(sitemap_url, loc))
elif kind == "urlset":
for item in root:
if local_name(item.tag) != "url":
continue
loc = child_text(item, "loc")
if not loc:
continue
row_key = (sitemap_url, loc)
if row_key in seen_pages:
continue
seen_pages.add(row_key)
page_rows.append({
"competitor_origin": origin,
"source_sitemap": sitemap_url,
"url_as_listed": loc,
"lastmod": child_text(item, "lastmod"),
"collected_at": started,
})
else:
errors.append({"url": sitemap_url, "error": f"Unexpected XML root: {kind}"})
except (requests.RequestException, ET.ParseError, OSError, ValueError) as exc:
errors.append({"url": sitemap_url, "error": str(exc)})
time.sleep(0.5) # Conservative pacing; tune only with the target's access guidance.
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
with open(f"urls-{stamp}.jsonl", "w", encoding="utf-8") as out:
for row in page_rows:
out.write(json.dumps(row, ensure_ascii=False) + "\n")
report = {"competitor_origin": origin, "started_at": started,
"sitemap_files": files, "url_count": len(page_rows), "errors": errors}
with open(f"report-{stamp}.json", "w", encoding="utf-8") as out:
json.dump(report, out, ensure_ascii=False, indent=2)
print(json.dumps({"url_count": len(page_rows), "files": len(files),
"errors": len(errors), "snapshot": stamp}, indent=2))
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("Usage: python collect_sitemaps.py https://www.example.com")
main(sys.argv[1])
For a production collector, consider adding a maximum response size, a persistent queue, conditional requests using ETag or Last-Modified response headers when available, and a database or object store. The script intentionally uses a simple local JSONL snapshot so its output can be reviewed and compared without a service dependency.
cURL: fetch and inspect a sitemap
cURL is useful for inspecting each location or automating a small known set. It does not itself recursively interpret XML indexes; use a parser for full extraction.
curl --fail --location --compressed \
--user-agent 'SitemapResearch/1.0 (+contact your-team@example.com)' \
--max-time 45 \
https://www.example.com/robots.txt
curl --fail --location --compressed \
--user-agent 'SitemapResearch/1.0 (+contact your-team@example.com)' \
--max-time 45 \
https://www.example.com/sitemap.xml -o sitemap.xml
Runnable Node.js index walker
Node.js 18 or later provides fetch. Save this as collect-sitemaps.mjs and run node collect-sitemaps.mjs https://www.example.com. It walks sitemap indexes and writes JSONL records; Node’s fetch transparently handles HTTP compression offered by servers, while the script also supports a gzip response body advertised as .gz.
import { writeFile } from 'node:fs/promises';
import { gunzipSync } from 'node:zlib';
const start = process.argv[2];
if (!start) throw new Error('Usage: node collect-sitemaps.mjs https://www.example.com');
const origin = new URL(start.includes('://') ? start : `https://${start}`).origin;
const ua = 'SitemapResearch/1.0 (+contact your-team@example.com)';
const timeoutMs = 45000;
const maxFiles = 10000;
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
const local = name => name.split('}').at(-1).toLowerCase();
const textOf = (parent, name) => {
for (const child of [...parent.children]) {
if (local(child.tagName) === name) return child.textContent.trim() || null;
}
return null;
};
async function fetchText(url) {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
const response = await fetch(url, { headers: { 'user-agent': ua }, signal: controller.signal,
redirect: 'follow' });
if (!response.ok) throw new Error(`HTTP ${response.status}`);
let bytes = Buffer.from(await response.arrayBuffer());
if (url.toLowerCase().split('?')[0].endsWith('.gz')) bytes = gunzipSync(bytes);
return { text: bytes.toString('utf8'), status: response.status };
} finally { clearTimeout(timer); }
}
function parseXml(xml) {
// Minimal XML tree parser using the platform DOMParser is not available in Node by default.
// Add a parser dependency for production use; this small extractor handles standard sitemap tags.
const decode = value => value.replaceAll('&', '&').replaceAll('<', '<')
.replaceAll('>', '>').replaceAll('"', '"').replaceAll(''', "'");
const rootMatch = xml.match(/<([\w:-]+)(?:\s[^>]*)?>/);
if (!rootMatch) throw new Error('No XML root element');
const root = local(rootMatch[1]);
const entries = [];
const blockName = root === 'sitemapindex' ? 'sitemap' : 'url';
const blocks = xml.match(new RegExp(`<${blockName}(?:\\s[^>]*)?>[\\s\\S]*?${blockName}\\s*>`, 'gi')) || [];
for (const block of blocks) {
const value = tag => {
const match = block.match(new RegExp(`<${tag}(?:\\s[^>]*)?>([\\s\\S]*?)${tag}\\s*>`, 'i'));
return match ? decode(match[1].trim().replace(//g, '$1')) : null;
};
entries.push({ loc: value('loc'), lastmod: value('lastmod') });
}
return { root, entries };
}
let queue = [];
const robotsUrl = `${origin}/robots.txt`;
try {
const { text } = await fetchText(robotsUrl);
queue = [...text.matchAll(/^\s*Sitemap\s*:\s*(\S+)/gim)].map(m => new URL(m[1], robotsUrl).href);
} catch (error) { console.error(`robots.txt: ${error.message}`); }
if (!queue.length) queue = ['/sitemap.xml', '/sitemap_index.xml', '/sitemap-index.xml'].map(p => `${origin}${p}`);
const seenFiles = new Set();
const seenPages = new Set();
const rows = [];
const errors = [];
const collectedAt = new Date().toISOString();
while (queue.length && seenFiles.size < maxFiles) {
const sitemap = queue.shift();
if (seenFiles.has(sitemap)) continue;
seenFiles.add(sitemap);
try {
const { text } = await fetchText(sitemap);
const { root, entries } = parseXml(text);
if (root === 'sitemapindex') {
for (const entry of entries) if (entry.loc) queue.push(new URL(entry.loc, sitemap).href);
} else if (root === 'urlset') {
for (const entry of entries) {
if (!entry.loc) continue;
const key = `${sitemap}\n${entry.loc}`;
if (seenPages.has(key)) continue;
seenPages.add(key);
rows.push({ competitor_origin: origin, source_sitemap: sitemap,
url_as_listed: entry.loc, lastmod: entry.lastmod, collected_at: collectedAt });
}
} else errors.push({ url: sitemap, error: `Unexpected XML root ${root}` });
} catch (error) { errors.push({ url: sitemap, error: error.message }); }
await sleep(500);
}
if (queue.length) errors.push({ error: `Stopped at MAX_FILES=${maxFiles}` });
const stamp = new Date().toISOString().replaceAll(':', '').replaceAll('-', '');
await writeFile(`urls-${stamp}.jsonl`, rows.map(row => JSON.stringify(row)).join('\n') + '\n');
await writeFile(`report-${stamp}.json`, JSON.stringify({ origin, collectedAt,
sitemap_files: [...seenFiles], url_count: rows.length, errors }, null, 2));
console.log(JSON.stringify({ url_count: rows.length, sitemap_files: seenFiles.size, errors: errors.length }));
The included Node parser is deliberately small and covers ordinary sitemap XML. If a target emits unusual XML constructs, use a maintained XML parser and validate against the sitemap namespace instead of expanding regular expressions.
3. Preserve provenance and compare snapshots
Store, at minimum, the competitor host, sitemap source, URL exactly as listed, collection timestamp, and lastmod if present. Keep raw values alongside any normalized comparison key. This lets you detect that a slash, hostname, case, query string, or redirect changed without silently rewriting the evidence.
Normalize consistently for comparison, and document the rules. A practical key may lowercase the hostname and remove its default port while preserving path case, query parameters, and trailing slash unless your analysis explicitly treats them as equivalent. Do not discard the original URL. Deduplicate exact repeats across child files, but record which source files listed each URL.
For two JSONL snapshots generated by the Python script, this small comparison identifies first-seen, disappeared, and retained raw URLs. Save as compare.py and run python compare.py older.jsonl newer.jsonl.
import json
import sys
def read(path):
with open(path, encoding='utf-8') as source:
return {json.loads(line)['url_as_listed'] for line in source if line.strip()}
if len(sys.argv) != 3:
raise SystemExit('Usage: python compare.py older.jsonl newer.jsonl')
old, new = read(sys.argv[1]), read(sys.argv[2])
for label, values in (('newly_listed', new - old), ('no_longer_listed', old - new),
('retained', old & new)):
print(f'[{label}] {len(values)}')
for url in sorted(values):
print(url)
A URL that appears is a new discovery event, not proof the page was just published. A URL that disappears is not proof the page was deleted. The sitemap may have been reorganized, temporarily incomplete, or corrected. Verify important changes with a fetch or crawl.
What does lastmod tell you?
lastmod represents the modification date of the linked page, not the time the sitemap was generated. The protocol accepts a date or a fuller W3C datetime. Google says it uses lastmod when the values are consistently and verifiably accurate. Keep it as publisher-provided metadata, record when you collected it, and compare it with observed content changes before relying on it as an alert trigger. Google ignores changefreq and priority for its crawling decisions; neither is evidence that a page changed.
4. Verify discoveries with a separate crawl
Sitemap extraction answers “what URLs did this sitemap expose on this date?” It does not answer whether those URLs work, are indexable, are linked internally, or contain a meaningful change. Fetch selected URLs separately and record:
- HTTP status and final URL after redirects
- Canonical link, if present
- Robots meta directives and
X-Robots-Tagresponse header - Whether robots rules allow the fetch under the access policy you are following
- Content type and a stable content fingerprint for page-level change detection
- Whether the URL is discoverable from internal links in your crawl
Compare sitemap URLs with crawl results to find pages listed but broken, blocked, canonicalized elsewhere, or absent from internal navigation. Sitebulb’s [sitemap source documentation](https://support.sitebulb.com/en/articles/9717059-website-crawl-sources) describes comparing sitemap sources with crawl data and identifying URL issues; this is an example of a relevant capability, not a recommendation for every budget or workflow.
For repeatable content comparison, fingerprint a normalized extraction of the main content rather than raw HTML alone. Raw HTML can change due to timestamps, ads, or rotating recommendations. Preserve the fetched timestamp and extraction method so a later comparison remains explainable.
5. Schedule collection and report coverage
Choose a schedule based on how quickly you need to notice changes and how frequently the site appears to update. There is no universally safe request rate in the sources used for this guide. Start conservatively, use modest concurrency, add backoff for transient errors, and review the site’s published access terms and instructions. For sensitive or high-stakes monitoring, get appropriate legal advice; public availability of a sitemap does not settle legal or contractual questions.
At scale, the workload depends on more than URL count: number of competitors, index depth, sitemap file count and size, fetch latency, parse errors, and how many discovered pages you subsequently crawl all matter. Track each file independently, retain HTTP status and collection time, and use validators such as ETag or Last-Modified where servers provide them. A failed child sitemap should be reported as partial coverage, not silently treated as an empty file.
Every run report should state the domains and subdomains in scope, discovered sitemap locations, sitemap files fetched, extraction timestamp, URL count, duplicate count, failed files, and any fallback paths tried. Use the same URL normalization and follow-up crawl settings between runs. If you publish alerts, label them “newly listed” or “no longer listed” until a page-level check confirms a publication, deletion, or content change.
6. Screenshot changed pages for visual review
When a page-level crawl flags a high-value URL, a screenshot can make visual changes easier to review alongside the URL diff. ScreenshotNeo is a website screenshot API and MCP server from [Yorker Media](https://screenshotneo.com). Its API can return PNG, JPEG, WebP, or PDF, and its parameter names are compatible with those used by other screenshot APIs to make switching easier. The examples below are for a page you are authorized to capture; use the service’s [API documentation](https://screenshotneo.com/docs/) for its options and response details.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Or skip the browser setup
Use ScreenshotNeo’s one-call API when you want a screenshot for a page in your monitoring set:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets can be removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the page verdict and billing status in headers. An MCP server gives AI agents tools to take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. See the [ScreenshotNeo API docs](https://screenshotneo.com/docs/) and [create a free account](https://screenshotneo.com/account/sign-up/).
7. Troubleshooting sitemap extraction
| Symptom | Likely cause | What to do |
|---|---|---|
robots.txt is missing or returns an error |
The host does not expose it, blocks the request, or is temporarily unavailable. | Record the status and timestamp. Try documented sitemap paths as labeled fallbacks; do not treat a failed robots fetch as proof no sitemap exists. |
| XML parser reports malformed content | The response may be an HTML error page, truncated, incorrectly encoded, or invalid XML. | Save the response status, content type, final URL, and a bounded sample for diagnosis. Confirm UTF-8 and inspect the XML before retrying. |
| A sitemap index yields no page URLs | It lists child sitemap files rather than page URLs. | Follow each <sitemap> location recursively and retain the parent-to-child chain. |
| Compressed file cannot be parsed | The response is gzip-compressed or the URL ends in .gz. |
Request with compression support and decompress gzip bytes before XML parsing. Avoid decompressing untrusted oversized content without a size limit. |
| Duplicate URLs appear | URLs may occur in multiple child sitemaps or repeat within a file. | Deduplicate exact values for counting while retaining all source sitemap references. |
lastmod stays unchanged or looks implausible |
The publisher may not maintain it accurately, or it may describe the sitemap file when read in an index context. | Preserve the value but verify important changes using fetched content and response metadata. |
| A listed URL redirects, fails, or canonicalizes elsewhere | The sitemap may be stale or may list an alias rather than the preferred page. | Record the listed URL and final URL separately; inspect status, canonical, and robots directives. |
| The URL count exceeds a parser or provider limit | The collection is split across child files, or a nonconforming oversized file is being served. | Process each child independently and report file-level failures. Google’s documented per-sitemap ceiling is 50 MB uncompressed or 50,000 URLs. |
| Two snapshots show thousands of changes at once | The sitemap was reorganized, a host changed, normalization changed, or a collection was partial. | Compare raw URLs and provenance, check redirects and failures, and rerun with the same scope before calling them page changes. |
8. Reliability, performance, and cost
Sitemap fetching is usually much cheaper than crawling every page because a run retrieves a finite set of sitemap documents first. Actual time and bandwidth depend on file sizes, response latency, and the number of children; the sources cited here provide protocol ceilings, not runtime benchmarks. Separate discovery from validation so you can fetch only changed or high-priority URLs instead of recrawling every listed page at every interval.
For reliability, checkpoint each completed file, retry transient failures with backoff, cap redirects and total files, set connection and read timeouts, and report partial completion. Retain raw snapshots so a later parser fix can be applied without fetching everything again. For larger collections, persist snapshots and diffs in a database or object store and alert on meaningful verified changes rather than raw sitemap churn.
The primary direct costs are compute, storage, bandwidth, and any crawler or screenshot services you choose to use. ScreenshotNeo’s published prices are Free for 1,000 shots per month with no card; Starter $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free. All features are included on every plan.
FAQ
Does a sitemap show every page on a website?
No. It shows URLs included in the sitemap files you successfully found and fetched. Pages may be omitted, and listed pages may not be live or indexed.
Should I treat a sitemap URL as proof a competitor published a new page?
No. Treat it as a newly observed URL. Confirm publication by checking the response and page content, and compare with prior crawl evidence where available.
Can I use only lastmod to detect content changes?
Use it as a clue, not the sole signal. It is useful when maintained accurately; corroborate important alerts with a page fetch or content fingerprint.
Do priority and changefreq help rank monitoring alerts?
They are publisher-provided fields, not observed evidence that a page changed. Google says it ignores both for crawling decisions.


