ScreenshotNeo

BlogHow-to

How to Download a Website Sitemap

Find a site's sitemap in robots.txt, download it with a browser or command line, follow sitemap indexes, and troubleshoot common errors.

By the ScreenshotNeo team1 October 20262 min read

Direct answer: Open https://example.com/robots.txt and find a line beginning with Sitemap:. Open that URL in a browser to inspect or save it, or download it from a terminal with curl -L -o sitemap.xml 'https://example.com/sitemap.xml'. If the response is a sitemap index, download the child sitemap URLs listed in its <loc> elements.

1. Find the sitemap URL

Replace example.com with the site’s exact host and open:

https://example.com/robots.txt

Look for one or more lines like:

Sitemap: https://www.example.com/sitemap.xml

A site can publish multiple sitemap URLs. Record every one. The Sitemap protocol allows sitemap locations to be declared in robots.txt. See Google’s robots.txt documentation.

If robots.txt has no Sitemap line

Try common paths, then check the site’s CMS documentation:

  • https://example.com/sitemap.xml
  • https://example.com/sitemap_index.xml
  • https://example.com/sitemap-index.xml

These are discovery guesses, not guaranteed locations. Many CMS platforms generate a sitemap automatically, but the filename and path vary.

2. Download it in a browser

  1. Paste the discovered sitemap URL into the address bar.
  2. Check that the response loads and looks like XML rather than an HTML error page.
  3. Use your browser’s save or download command to store the response locally.

This is the simplest method for a one-off check because you can immediately see whether the URL responds. Google also recommends opening a sitemap URL in a browser when troubleshooting.

3. Download it from the command line

cURL

curl -L -o sitemap.xml 'https://www.example.com/sitemap.xml'

-L follows redirects and -o writes the response to a file. Add headers when a site requires them:

curl -L \
  -A 'Mozilla/5.0' \
  -H 'Accept: application/xml,text/xml;q=0.9,*/*;q=0.8' \
  -o sitemap.xml \
  'https://www.example.com/sitemap.xml'

To inspect status and response headers without saving the body:

curl -I -L 'https://www.example.com/sitemap.xml'

wget

wget -O sitemap.xml 'https://www.example.com/sitemap.xml'

Use the exact URL found in robots.txt; do not assume the extension alone proves that the response is a sitemap.

Python

import requests

url = "https://www.example.com/sitemap.xml"
response = requests.get(url, timeout=30)
response.raise_for_status()
with open("sitemap.xml", "wb") as file:
    file.write(response.content)
print(f"Saved {len(response.content)} bytes to sitemap.xml")

Node.js

import { writeFile } from "node:fs/promises";

const url = "https://www.example.com/sitemap.xml";
const response = await fetch(url, { redirect: "follow" });
if (!response.ok) {
  throw new Error(`HTTP ${response.status} ${response.statusText}`);
}
const body = Buffer.from(await response.arrayBuffer());
await writeFile("sitemap.xml", body);
console.log(`Saved ${body.length} bytes to sitemap.xml`);

4. Tell a sitemap index from a URL sitemap

Inspect the root element after downloading. A <sitemapindex> contains links to other sitemap files:

<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemap-posts.xml</loc>
  </sitemap>
</sitemapindex>

Download each child <loc> URL. A <urlset> directly lists pages:

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/page/

The filename is not authoritative. Inspect the XML response and root element. For the supported formats and protocol details, see Google Search Central's sitemap documentation.

Follow an index with Python

from urllib.parse import urljoin
import requests
import xml.etree.ElementTree as ET

INDEX_URL = "https://www.example.com/sitemap_index.xml"
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}

xml = requests.get(INDEX_URL, timeout=30)
xml.raise_for_status()
root = ET.fromstring(xml.content)

if root.tag.endswith("sitemapindex"):
    child_urls = [node.text.strip() for node in root.findall("sm:sitemap/sm:loc", NS)]
    for number, child_url in enumerate(child_urls, start=1):
        child = requests.get(child_url, timeout=30)
        child.raise_for_status()
        with open(f"sitemap-{number}.xml", "wb") as file:
            file.write(child.content)
else:
    with open("sitemap.xml", "wb") as file:
        file.write(xml.content)

print("Downloaded sitemap files")

5. Validate and inspect the downloaded file

  • Confirm the HTTP status is successful and the body is XML.
  • Check whether the root is urlset or sitemapindex.
  • Read each loc value and verify its scheme and host.
  • Watch for an HTML login page, access-denied page, or proxy error saved with an .xml filename.
  • For compressed files such as sitemap.xml.gz, decompress them before parsing.

Google documents a limit of 50,000 URLs or 50 MB uncompressed per sitemap file. Larger sites split URLs across several files and expose them through an index.

6. Download every child sitemap

For a small index, extract the <loc> values and download them one at a time. For a repeatable workflow, parse the XML and add retries, a timeout, and a bounded request rate. Preserve the original URLs because sitemap indexes can use absolute URLs on a different host or path.

7. Troubleshooting

Symptom Likely cause Fix
404 Not Found Wrong path, host, or protocol Recheck the domain and the Sitemap line in robots.txt; try the site's documented CMS path.
403 or 429 Access controls or rate limiting Slow requests, use the canonical URL, check whether authentication is required, and retry later.
Downloaded file contains HTML Error page, redirect target, login page, or bot challenge Inspect status and headers with curl -I -L, then open the final URL in a browser.
XML parse error Truncated download, invalid encoding, or non-XML response Redownload, compare file size, inspect the first bytes, and validate the XML.
Only an index appears The index intentionally points to child files Follow every child <loc> URL and inspect each resulting sitemap.
Browser shows a blank page Browser XML rendering behavior or an empty response Save the response with cURL and inspect its bytes and HTTP status.
Google does not process it Invalid format, blocked access, or size limits Check Search Console's Sitemaps report, validate the file, and split files that exceed documented limits.

A sitemap discovered outside Search Console may not appear in its Sitemaps report. Submission is a hint; it does not guarantee that Google will fetch, crawl, or index every listed URL.

8. Browser versus command line

Method Best for Trade-off
Browser One-off inspection and quick troubleshooting Manual saving and repetition
cURL or wget Saving exact responses and scripting You must inspect status and content yourself
Python or Node.js Following indexes and integrating with pipelines Requires parsing, retries, and error handling

9. Performance, reliability, and cost notes

  • Download child files sequentially or with a small concurrency limit so you do not trigger rate limits.
  • Set connection and read timeouts; retry transient 5xx responses with backoff.
  • Cache files when auditing repeatedly, and record retrieval time and HTTP status.
  • Stream very large responses to disk instead of holding them all in memory.
  • Downloading a sitemap is normally inexpensive because it is a direct HTTP request; crawler software is only needed for broader site analysis.
  • A successful HTTP response does not prove the XML is valid, complete, or useful for indexing.

10. Or skip the browser setup

If you need a visual snapshot of a sitemap or its rendered page, ScreenshotNeo can capture the URL with one request. It accepts a URL and returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.example.com/sitemap.xml -o sitemap.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.example.com/sitemap.xml"}, timeout=90)
open("sitemap.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.example.com/sitemap.xml' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Where is a website's sitemap usually located?

Start with /robots.txt. If it is not declared there, try common CMS paths such as /sitemap.xml and /sitemap_index.xml.

Can I download a sitemap without Google Search Console?

Yes. A browser, cURL, wget, Python, or Node.js can request the sitemap URL directly.

What if the sitemap has more than 50,000 URLs?

It should be split into multiple sitemap files, usually referenced by a sitemap index. Google documents both the 50,000 URL and 50 MB uncompressed limits.

Does downloading a sitemap make pages get indexed?

No. A sitemap helps search engines discover URLs, but submission is only a hint and does not guarantee crawling or indexing.