ScreenshotNeo

BlogAI agents

How to Capture Screenshots of Pages from an Indian Sitemap with MCP

Parse sitemap URLs, capture pages with Playwright MCP, and save auditable screenshots. Includes full-page options, failure handling, and an API shortcut.

By the ScreenshotNeo team4 October 202610 min read

A sitemap gives you candidate page URLs; an MCP browser can visit those absolute URLs and save screenshots. For a sitemap index, first fetch its child sitemaps, then extract page URLs from each one. With Playwright MCP, navigate to each URL and choose a viewport, element, or full-page screenshot. There is no special India-specific sitemap format implied here: without a specific domain, this is a general workflow, not a walkthrough verified against an Indian site.

1. Find the sitemap URL

Start with the website’s published sitemap location or its robots.txt file, which may declare a Sitemap: URL. Do not assume the sitemap lives at /sitemap.xml; no universal path is guaranteed. Use the exact location declared by the site or provided by its owner. Google’s sitemap guidance covers sitemap locations and URL requirements.

A sitemap is a source of URLs, not proof that each page is public, loads in an automated browser, or is indexed. Google describes submission as a hint rather than a guarantee of crawling. A browser capture is a separate operation with its own access and loading conditions.

2. Determine whether it is a sitemap index

Fetch the XML and inspect its root element. A urlset contains page entries, usually with a loc child for each page. A sitemapindex contains child sitemap entries; fetch those child locations and extract their page URLs in turn. Google documents the sitemap index structure and limits, including up to 50,000 loc entries in an index.

Sitemap files should be UTF-8, and listed page URLs should be fully qualified absolute URLs. Use each listed URL as supplied; do not prepend the domain to a relative path unless you are deliberately correcting invalid source data.

3. Set up Playwright MCP

Use an MCP client connected to a Playwright MCP server that exposes browser navigation and screenshot tools. The relevant documented operations are browser_navigate and browser_take_screenshot. Tool calls are issued through your MCP client; exact setup and call syntax can vary by client and server version, so follow that client’s MCP configuration and tool-call interface.

Playwright MCP supports structured accessibility snapshots for inspecting page structure and locating elements, and screenshots for visual verification. Consult the official Playwright MCP project documentation for current server setup and tool schemas.

4. Parse sitemap XML and prepare a URL list

The following runnable Python script fetches a sitemap or sitemap index, follows child sitemap locations recursively, handles gzip-compressed sitemap responses, and prints unique page URLs. Install the dependency with python -m pip install requests. Run it with a sitemap URL, for example python sitemap_urls.py https://example.in/sitemap.xml; replace that example with the actual declared sitemap URL.

import gzip
import sys
import requests
import xml.etree.ElementTree as ET


def local_name(tag):
    return tag.rsplit("}", 1)[-1]


def fetch_xml(url):
    response = requests.get(url, timeout=30)
    response.raise_for_status()
    data = response.content
    if url.endswith(".gz") or response.headers.get("Content-Type", "").lower().endswith("gzip"):
        try:
            data = gzip.decompress(data)
        except OSError:
            pass
    return ET.fromstring(data)


def loc_values(root, parent_name):
    values = []
    for parent in root:
        if local_name(parent.tag) != parent_name:
            continue
        for child in parent:
            if local_name(child.tag) == "loc" and child.text:
                values.append(child.text.strip())
    return values


def collect_pages(sitemap_url, seen_sitemaps=None):
    if seen_sitemaps is None:
        seen_sitemaps = set()
    if sitemap_url in seen_sitemaps:
        return []
    seen_sitemaps.add(sitemap_url)
    root = fetch_xml(sitemap_url)
    kind = local_name(root.tag)
    if kind == "urlset":
        return loc_values(root, "url")
    if kind == "sitemapindex":
        pages = []
        for child_sitemap in loc_values(root, "sitemap"):
            pages.extend(collect_pages(child_sitemap, seen_sitemaps))
        return pages
    raise ValueError(f"Unexpected XML root: {kind}")


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python sitemap_urls.py SITEMAP_URL")
    for page_url in dict.fromkeys(collect_pages(sys.argv[1])):
        print(page_url)

This simple script follows the XML structure and avoids revisiting the same child sitemap. For a very large sitemap collection, stream or persist results instead of holding every URL in memory. Treat parsed URLs as untrusted input: restrict captures to the intended site and avoid passing arbitrary URLs from an untrusted sitemap into a browser running with access to sensitive internal networks.

5. Capture each page with MCP

  1. Take the absolute page URL from the parsed sitemap output.
  2. Call Playwright MCP’s navigation tool for that URL.
  3. If you need a particular component, inspect the page with an accessibility snapshot and identify its locator or selector.
  4. Call the screenshot tool with an output path and the desired scope.
  5. Record the source URL alongside the saved filename so each image can be traced back to the sitemap entry.

The following illustrates the tool-call sequence in pseudocode because MCP clients represent tool invocations differently. Use the actual argument names shown by your connected Playwright MCP server.

for url in sitemap_page_urls:
    await browser_navigate({"url": url})
    # Optional: inspect with the client's accessibility snapshot tool.
    await browser_take_screenshot({
        "filename": filename_for(url),
        "fullPage": True
    })

Playwright’s screenshot documentation describes capturing the viewport, a specific element, or the full scrollable page. See the official Playwright MCP screenshots documentation for available arguments and output behavior in your installed version.

6. Choose screenshot scope and output format

Need Capture Notes
Record what is visible now Viewport screenshot Useful for consistent, fixed-size captures; content below the fold is excluded.
Show one chart or page region Element screenshot Pass a target element or selector supported by the tool. Inspect the page first if you do not know the selector.
Include below-the-fold content Full-page screenshot Uses fullPage: true. It cannot be combined with a target element in the documented workflow.

Supported screenshot formats include PNG, JPEG, and WebP. A filename extension can determine the format when a type is not explicitly supplied. Choose a format based on downstream use: PNG preserves sharp text and UI edges well; JPEG and WebP can reduce file size depending on content and encoding settings. Verify the actual tool schema and output path behavior in your MCP server version.

Lazy-loaded images and content that appears after scrolling may not be present until the page has had time to render or been scrolled. If those sections matter, wait for a known selector or explicitly scroll through the page before capturing, where the browser tool supports those actions.

7. Name files and keep an audit record

Use stable filenames derived from a normalized page path, and keep a manifest containing the original URL, output path, capture time, and result. This naming convention is practical guidance, not a Playwright requirement. Avoid using the full URL as a raw filename because query characters and long paths can be awkward across operating systems.

https://example.in/guides/payments/  ->  example-in__guides__payments.png

If multiple sitemap entries normalize to the same filename, add a short stable hash or numeric suffix. Preserve query strings in the manifest even if the filename omits them.

8. cURL, Python, and Node.js alternatives for fetching the sitemap

MCP performs the browser navigation and screenshot action. These snippets only retrieve a sitemap response; they do not make a browser screenshot. They can help inspect the XML or feed URLs to an MCP client.

cURL

curl --fail --location --compressed \
  --output sitemap.xml \
  "https://example.in/sitemap.xml"

Python

import requests

url = "https://example.in/sitemap.xml"
response = requests.get(url, timeout=30)
response.raise_for_status()
with open("sitemap.xml", "wb") as output:
    output.write(response.content)

Node.js

const response = await fetch('https://example.in/sitemap.xml');
if (!response.ok) {
  throw new Error(`Sitemap request failed: ${response.status}`);
}
const xml = await response.text();
await import('node:fs/promises').then(fs => fs.writeFile('sitemap.xml', xml));

For a sitemap index, repeat retrieval for each child sitemap URL and parse its XML. A raw text search for <loc> is fragile because XML namespaces, escaping, and formatting can vary; use an XML parser for production workflows.

9. Reliability, speed, and cost considerations

  • Bound concurrency. A large sitemap can contain many pages. Capture in small batches and limit simultaneous browser pages to avoid resource exhaustion and unnecessary load on the site.
  • Use explicit readiness conditions. A navigation event alone may precede client-rendered content. Wait for a meaningful page selector or a bounded delay when required; avoid indefinite network-idle waits on pages with continuous analytics or streaming requests.
  • Retry selectively. Retry transient network errors with a limit and backoff. Do not retry persistent 404, access-denied, or malformed-URL failures blindly.
  • Make runs resumable. Record completed URLs and output files so a process restart does not repeat every capture.
  • Expect variability. Cookie dialogs, bot challenges, geolocation, authentication, and dynamic content can make screenshots differ or prevent capture. A sitemap listing does not guarantee browser access.
  • Budget storage and runtime. Full-page images and high-resolution pages take more time and disk space than viewport images. Estimate from a small representative sample before running a large collection.
  • Respect site access rules. A sitemap is not permission to overload a site or bypass access controls. Keep request rates modest and use authorized credentials only.

Google’s robots guidance is about crawler access and search visibility. A disallowed URL can still appear in search results under some conditions; robots.txt should not be treated as a guaranteed removal mechanism. Those search crawler rules do not by themselves establish whether your browser automation can reach a page. See Google’s robots.txt guide.

10. Troubleshooting

Symptom Likely cause Fix
404 when requesting /sitemap.xml The site uses a different sitemap location or does not publish one there. Check the sitemap declaration in robots.txt or request the correct URL from the site owner.
XML parser reports an unexpected root The response is HTML, an error page, a sitemap format you did not handle, or a compressed file decoded incorrectly. Check status code, content type, response body, and root element. Handle urlset and sitemapindex explicitly.
No URLs found Wrong namespace handling, empty sitemap, or parsing only index children as page entries. Match elements by local XML name and branch on sitemapindex versus urlset.
Browser navigation times out Slow origin, stalled requests, bot check, or an overly strict readiness condition. Inspect the page state, increase a bounded timeout when reasonable, wait for a specific selector, and mark unresolved URLs for review.
Screenshot is blank or incomplete Capture occurred before rendering, content is lazy-loaded, or the page returned an error/challenge. Inspect with an accessibility snapshot and browser state, wait for content, scroll if necessary, then capture again if the failure is transient.
Full-page and element options conflict The documented screenshot mode does not combine a target with full-page capture. Take separate screenshots: one full page and one element capture.
Files overwrite each other Different URLs map to the same slug. Add a deterministic hash or unique suffix and retain the exact URL in a manifest.
Many captures fail at once Excess concurrency, rate limits, session state, or a shared network problem. Reduce concurrency, check connectivity and server responses, and resume only failed entries after recovery.

11. Or skip the browser setup

If the goal is a clean image per URL and you do not need an MCP-controlled browser, ScreenshotNeo is a website screenshot API and MCP server. The API can return PNG, JPEG, WebP, or PDF from one GET request. Put each sitemap URL into the url parameter and save the response using the chosen output format. See the ScreenshotNeo API documentation for parameters and response details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.in/ \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.in/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.in/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(async fs =>
  fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

12. FAQ

Does an Indian sitemap use a different XML format?

The country reference alone does not imply a different format. Use the sitemap or index structure published by the specific site.

Does a sitemap URL guarantee the page can be captured?

No. The URL is a candidate to visit. It can redirect, require authentication, fail, or show a bot challenge.

Can I take both full-page and element screenshots?

Yes, as separate captures. The documented full-page option cannot be combined with a target element.

Does robots.txt determine whether Playwright MCP can navigate?

Robots.txt communicates crawler rules. It should not be conflated with browser network access or authorization.

Can I capture every URL in a large sitemap at once?

You can automate a collection, but use bounded concurrency, resumable records, and a rate appropriate for the site.