ScreenshotNeo

BlogGuides

Website Screenshot APIs That Capture Pages from a Sitemap

Capture sitemap URLs with a native batch API or extract them and automate individual screenshots. Compare the workflows, handle sitemap indexes, and avoid missed pages.

By the ScreenshotNeo team4 October 202611 min read

Yes. You can capture pages listed in a sitemap with an API that supports sitemap discovery, or fetch the sitemap yourself and submit each URL to a screenshot API. A native batch endpoint can discover sitemap indexes, capture pages, and package results in one asynchronous job. The DIY route gives you more control over URL filtering, retries, concurrency, and output naming, but you must build those pieces.

A sitemap is a discovery list, not a guarantee that every page is present or capturable. Pages can be missing from it, inaccessible, redirected, blocked, or fail while rendering. Choose the workflow that matches whether you need sitemap discovery, individual image control, or an archive of many captures.

1. Choose a sitemap screenshot workflow

Workflow URL discovery Output and job handling Good fit
ScreenshotNeo — clean screenshots, only clean shots billed, and a paid plan starting at $5 for 3,000 shots Fetch and parse the sitemap yourself, then submit each URL One GET request per page; automate the list and save each image Individual PNG, JPEG, or WebP captures; custom capture options; scripts and AI-agent workflows
EnConvert website-to-screenshot endpoint Native sitemap mode; fetches /sitemap.xml and recursively follows sitemap indexes Asynchronous batch; returns HTTP 202 and a batch ID, then a status endpoint or notification provides the ZIP when ready A managed sitemap-to-ZIP workflow
Urlbox with a sitemap extractor or CaptureDeck Extract URLs first; Urlbox’s guide describes this as a two-step workflow, not native whole-site sitemap capture in one request Submit extracted URLs through automation or the no-code CaptureDeck option Teams already using Urlbox or preferring a no-code workflow

EnConvert documents that sitemap mode requires a paid plan and private API key, and its maximum pages depend on plan limits. Its endpoint captures full-page PNGs and packages them into a ZIP. Urlbox documents extracting the URL list first, then using CaptureDeck or custom API automation. These are vendor-documented workflows, not an independent performance or pricing comparison. EnConvert documentation; Urlbox guide.

2. Use a native sitemap batch endpoint

EnConvert documents POST /v1/convert/website-to-screenshot. Set crawl_mode to sitemap to use sitemap discovery explicitly. The default auto mode uses the highest crawl mode allowed by the account plan; that may be full crawling rather than sitemap-only. A sitemap job fetches {base_url}/sitemap.xml; if it finds a sitemap index, it fetches child sitemaps and collects URL entries. The response is asynchronous: submit, save the batch ID, poll the status endpoint, and retrieve the ZIP URL when complete.

curl -X POST "https://api.enconvert.com/v1/convert/website-to-screenshot" \
  -H "Content-Type: application/json" \
  -H "X-API-Key: $ENCONVERT_PRIVATE_KEY" \
  -d '{"url":"https://example.com","crawl_mode":"sitemap"}'

The response contains a batch ID. Poll /v1/convert/batch/{batch_id} with the same private key until the job status is completed, partial, or failed. The status response can include a presigned ZIP download URL. Treat the returned page count and final status as job output; do not assume every discovered URL produced an image.

curl "https://api.enconvert.com/v1/convert/batch/$BATCH_ID" \
  -H "X-API-Key: $ENCONVERT_PRIVATE_KEY"

For production automation, use the documented webhook or email notification if polling is inconvenient. Handle partial and failed statuses explicitly, persist the batch ID, and download the archive before its presigned URL expires. Verify current endpoint details and plan limits in the EnConvert endpoint documentation.

3. DIY: parse a sitemap and capture each URL

For APIs without native sitemap discovery, the workflow is: find the sitemap, recursively parse sitemap indexes, normalize and filter URLs, then submit each page to the screenshot endpoint. The example below uses Python’s standard XML parser and ScreenshotNeo’s screenshot API. It saves one WebP per discovered URL. The script includes a timeout, a small concurrency limit, retries for transient failures, and safe filenames.

Find the sitemap first

Check https://example.com/robots.txt for a Sitemap: directive. Common paths include /sitemap.xml, /sitemap_index.xml, and /sitemaps.xml. A site may have no public sitemap, may publish more than one, or may require an alternate host. A sitemap index contains links to other sitemap documents; a URL set contains page URLs in <loc> elements.

Runnable Python script

Install the two dependencies with python -m pip install requests. Set SCREENSHOTNEO_API_KEY in the environment. Keep API keys out of source control and client-side code.

import concurrent.futures
import hashlib
import os
import time
import xml.etree.ElementTree as ET
from urllib.parse import urlparse

import requests

ROOT = "https://example.com"
API = "https://api.screenshotneo.com/v1/shot"
KEY = os.environ["SCREENSHOTNEO_API_KEY"]
OUT = "screenshots"
MAX_SITEMAPS = 100
WORKERS = 3

session = requests.Session()
session.headers.update({"User-Agent": "SitemapScreenshot/1.0"})

def local_name(tag):
    return tag.rsplit("}", 1)[-1]

def fetch_xml(url):
    response = session.get(url, timeout=(10, 30))
    response.raise_for_status()
    return ET.fromstring(response.content)

def discover(start_url):
    pending = [start_url]
    seen_sitemaps = set()
    pages = []
    while pending:
        sitemap_url = pending.pop()
        if sitemap_url in seen_sitemaps:
            continue
        if len(seen_sitemaps) >= MAX_SITEMAPS:
            raise RuntimeError("Sitemap limit reached; inspect sitemap index and raise MAX_SITEMAPS if expected")
        seen_sitemaps.add(sitemap_url)
        root = fetch_xml(sitemap_url)
        kind = local_name(root.tag)
        if kind == "sitemapindex":
            for node in root.iter():
                if local_name(node.tag) == "loc" and node.text:
                    pending.append(node.text.strip())
        elif kind == "urlset":
            for node in root.iter():
                if local_name(node.tag) == "loc" and node.text:
                    pages.append(node.text.strip())
        else:
            raise ValueError(f"Unexpected sitemap XML root: {kind}")
    # Preserve order while removing duplicate URLs.
    return list(dict.fromkeys(pages))

def safe_name(url):
    parsed = urlparse(url)
    label = (parsed.netloc + parsed.path).strip("/") or "home"
    label = "".join(c if c.isalnum() or c in "-_." else "_" for c in label)
    return f"{label[:120]}-{hashlib.sha256(url.encode()).hexdigest()[:10]}.webp"

def capture(url):
    params = {"access_key": KEY, "url": url}
    for attempt in range(4):
        try:
            response = session.get(API, params=params, timeout=(10, 90))
            response.raise_for_status()
            # Save the response body only after HTTP success. Check response headers
            # or your account's API behavior if you need to classify page verdicts.
            path = os.path.join(OUT, safe_name(url))
            with open(path, "wb") as image_file:
                image_file.write(response.content)
            return (url, "saved", path)
        except requests.RequestException as exc:
            if attempt == 3:
                return (url, "failed", str(exc))
            time.sleep(2 ** attempt)

os.makedirs(OUT, exist_ok=True)
# Prefer the sitemap URL advertised in robots.txt if the site uses a nonstandard path.
sitemap_url = ROOT.rstrip("/") + "/sitemap.xml"
urls = discover(sitemap_url)
print(f"Discovered {len(urls)} unique URLs")
with concurrent.futures.ThreadPoolExecutor(max_workers=WORKERS) as pool:
    for result in pool.map(capture, urls):
        print(*result)

Set ROOT to the site origin and adjust sitemap_url if robots.txt advertises a different sitemap. The script deliberately bounds workers. If your sitemap has many URLs, save progress to a manifest or database so a process restart does not recapture completed pages. Review the site’s terms and crawling policy and keep request rates reasonable.

cURL: capture one discovered URL

After extracting a page URL, pass it as a URL-encoded query parameter. Repeat this call for each sitemap entry in your orchestration script.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/docs \
  -o docs.webp

Python: capture one URL

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/docs"},
    timeout=90,
)
r.raise_for_status()
open("docs.webp", "wb").write(r.content)

Node.js: capture one URL

This uses built-in fetch in a current Node.js release. Check HTTP status before writing the response body.

import { writeFile } from 'node:fs/promises';

const q = new URLSearchParams({
  access_key: process.env.SCREENSHOTNEO_API_KEY,
  url: 'https://example.com/docs'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`, {
  signal: AbortSignal.timeout(90000)
});
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('docs.webp', Buffer.from(await res.arrayBuffer()));

For ScreenshotNeo capture parameters, response details, and options, see the ScreenshotNeo API documentation.

4. Or skip the browser setup

Fetch and parse the sitemap yourself, then call ScreenshotNeo once per page; it does not discover or batch sitemap URLs for you. One GET request returns an image or PDF for a URL, and its API supports PNG, JPEG, or WebP output.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/docs \
  -o shot.webp

See the API docs for request options. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card required.

5. Sitemap details that affect completeness

  • Sitemap indexes: A root sitemap may list child sitemap files rather than pages. Follow each child recursively and deduplicate URLs.
  • Namespaces: XML often uses a namespace. Parse elements by local tag name or use the correct namespace mapping; a plain unqualified .findall("url/loc") can return no matches.
  • Compressed sitemaps: Some sites publish gzip files such as sitemap.xml.gz. Decompress the response before XML parsing. Confirm the response is actually XML or gzip data rather than an HTML error page.
  • Multiple sitemap locations: Read robots.txt, check the site’s canonical host, and consider alternate locale or product sitemaps when the owner publishes them.
  • URL scope: Sitemaps may list multiple hosts or URL variants. Decide whether to capture every listed origin, only the target host, or a defined path subset.
  • Canonicalization and duplicates: Deduplicate exact URLs first. Do not aggressively remove query strings or normalize case without understanding the site’s routing; query parameters may identify distinct pages.
  • Not all pages are listed: A sitemap can be stale or incomplete. If you need pages discoverable only through links, a link crawler is a separate discovery method and needs scope controls.
  • Access and rendering: A sitemap can list pages that redirect, require authentication, reject automated requests, render slowly, or depend on client-side state. Discovery does not guarantee a successful screenshot.

6. Batch design, reliability, and cost

Control concurrency and protect the target

Start with a small number of concurrent captures, then adjust based on the target site’s response behavior and your API’s documented limits. Avoid launching an unbounded request per URL: it can overload the target, trigger rate limits, or exhaust local sockets and memory. Add a delay or token bucket when the site asks for slower access. Keep crawl and capture concurrency configurable rather than hard-coded in a production job.

Make runs resumable

  1. Persist the discovered URL list and a stable identifier for each URL.
  2. Record status, attempt count, last error, and output path for each page.
  3. Retry transient network failures and server errors with exponential backoff and jitter; do not retry permanent invalid-URL or authorization errors indefinitely.
  4. Use idempotent filenames or a manifest so reruns do not silently overwrite a different page.
  5. Verify the output is a valid image before marking the URL complete, and keep failures in a report for targeted reruns.

For asynchronous native APIs, persist the batch ID and poll at a measured interval or use a documented callback. Handle completed, partial, and failed states. Treat download URLs as temporary if the provider says they are presigned.

Plan for output storage and spend

Total work grows with the number of discovered URLs and the cost per successful capture. Estimate page count before a large run, account for retries, and decide whether you need full-page captures or a smaller viewport. Store only the formats and versions you need, and expire temporary archives or intermediate files. Check a provider’s current plan caps, batch size, rate limits, storage window, and authentication options before scheduling a recurring run.

ScreenshotNeo bills only clean shots: bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Its listed plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan. Match a recurring sitemap job’s expected successful captures to the plan allowance, and account for reruns and pages added to the sitemap.

7. Troubleshooting sitemap capture

Symptom Likely cause Fix
Native sitemap job says sitemap is missing The provider checks a standard path but the site publishes a different location, or has no public sitemap Inspect robots.txt and the site’s sitemap links. If the endpoint accepts only a base URL and fixed path, extract the sitemap yourself and submit URLs individually.
Non-200 sitemap response or timeout Server/CDN issue, access restriction, slow response, or unsupported host routing Request the sitemap directly, check status and redirects, and retry later for transient failures. Confirm the endpoint’s fetch timeout and supported authentication.
Invalid XML or zero URLs HTML error/challenge returned as XML, malformed document, unexpected namespace, empty sitemap, or parser handles only URL sets Inspect the response body and content type; parse namespaces; support both sitemapindex and urlset; reject unexpected document roots clearly.
Only some pages appear in the archive Some URLs failed during rendering, the sitemap was incomplete, or the batch ended partially Inspect per-page or batch status, capture failed URLs separately, and compare the discovered URL list to the output manifest.
401 or 403 from the screenshot API Missing, invalid, or wrong type of API key; some batch endpoints require a private key Use the credential type and header/query parameter required by that API. Keep secrets in environment variables and check account access.
429 or target-site blocks Rate limits or bot defenses Reduce concurrency, add spacing and backoff, and confirm you are permitted to capture the pages. Do not retry rapidly in a tight loop.
Saved file is HTML or an error payload Code wrote a non-image response without checking status or content Check HTTP status and response headers before saving. Validate file signatures or decode the image, then record the response as failed if invalid.
Duplicate or overwritten filenames Path-only naming collides for query URLs or repeated slugs Include a stable hash of the full URL, as in the Python example, and retain the original URL-to-file manifest.
Capture is blank or misses lazy content Page requires more render time, scrolling, or interaction; content may require login or be blocked Use wait conditions, full-page capture, or relevant interaction options offered by the screenshot API. Verify the page is publicly accessible and compare the rendered URL in a browser.

8. FAQ

Does a sitemap guarantee every page on the site is captured?

No. It only supplies listed URLs, and each listed page can still fail, redirect, or require access the capture process does not have.

Can I use ScreenshotNeo to submit a sitemap and get one ZIP?

ScreenshotNeo’s documented pattern here is one request per page URL. Parse the sitemap in your own script and orchestrate those requests; the API call does not discover sitemap URLs or create a sitemap batch archive.

Should I use sitemap mode or a full crawl?

Use sitemap discovery when the sitemap is the intended scope. A link crawl can discover unlisted pages, but it needs explicit domain, path, depth, and URL-trap controls.

Can I capture pages behind login?

Only if the selected service and endpoint support the required authentication mechanism and you provide valid credentials or session state. Check the relevant API documentation and plan restrictions before relying on it.