How to Capture Scheduled Screenshots of Every Page in a Sitemap
Build a scheduled sitemap screenshot workflow with Playwright, GitHub Actions, and reliable storage, then automate captures with ScreenshotNeo.
To capture scheduled screenshots of every page in a sitemap, parse the sitemap (including any child sitemaps), deduplicate its URLs, visit each page with a browser such as Playwright, save a full-page screenshot and a manifest, then run the script on a schedule. A sitemap is an inventory of the URLs it lists, not proof that it includes every route or interactive state on a site.
This guide uses Playwright with Python and GitHub Actions. It also covers local cron, URL scope, rate limits, visual comparisons, failures, and a managed API option. Confirm you are authorized to capture the target site and follow applicable robots.txt directives.
1. Decide what “every page” means
Write down the capture scope before automating. Decide which sitemap URL or sitemap index to use, which hosts are allowed, whether query strings matter, and how to handle locale variants, redirects, authenticated pages, and URL fragments. A sitemap pass covers sitemap-listed URLs within that scope. It does not exercise forms, menus, or routes absent from the sitemap.
The Sitemaps protocol permits up to 50,000 URLs and 50 MB in a sitemap file, and supports sitemap index files that point to child sitemaps. Plan to process child files when the starting URL is an index. See the Sitemaps protocol.
| Decision | Practical default |
|---|---|
| Allowed hosts | Restrict captures to the site and approved subdomains. |
| Query parameters | Keep only parameters that represent distinct pages; remove tracking parameters before deduplication. |
| Fragments | Drop URL fragments for page-level screenshots; capture separate states explicitly if needed. |
| Locale variants | Keep each locale URL if each is a distinct page to review. |
| Authentication | Use a dedicated test account and protected secrets; do not put credentials in source control. |
| History | Choose a retention period and archive layout before runs accumulate. |
2. Create a Playwright sitemap capture script
The script below reads a sitemap or sitemap index, fetches child sitemap files, normalizes and deduplicates URLs, applies a host allowlist, captures full pages, and writes a JSONL manifest. It handles gzip sitemap files, XML namespaces, redirects, and per-page errors. Install its dependencies with:
python -m pip install playwright requests
python -m playwright install chromium
Save this as capture_sitemap.py:
import gzip
import hashlib
import json
import os
import re
import time
import xml.etree.ElementTree as ET
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import parse_qsl, urlencode, urljoin, urlsplit, urlunsplit
import requests
from playwright.sync_api import sync_playwright
SITEMAP_URL = os.environ.get("SITEMAP_URL", "https://example.com/sitemap.xml")
OUTPUT_DIR = Path(os.environ.get("OUTPUT_DIR", "captures"))
ALLOWED_HOSTS = {
host.strip().lower()
for host in os.environ.get("ALLOWED_HOSTS", "example.com,www.example.com").split(",")
if host.strip()
}
MAX_PAGES = int(os.environ.get("MAX_PAGES", "50000"))
DELAY_SECONDS = float(os.environ.get("DELAY_SECONDS", "1.0"))
NAVIGATION_TIMEOUT_MS = int(os.environ.get("NAVIGATION_TIMEOUT_MS", "45000"))
READY_SELECTOR = os.environ.get("READY_SELECTOR", "")
session = requests.Session()
session.headers.update({"User-Agent": "SitemapScreenshotJob/1.0"})
def fetch_sitemap(url):
response = session.get(url, timeout=30)
response.raise_for_status()
body = response.content
if urlsplit(response.url).path.endswith(".gz") or body[:2] == b"\\x1f\\x8b":
body = gzip.decompress(body)
root = ET.fromstring(body)
kind = root.tag.rsplit("}", 1)[-1]
locations = [element.text.strip() for element in root.iter() if element.tag.rsplit("}", 1)[-1] == "loc" and element.text]
return kind, locations
def canonicalize(url):
parts = urlsplit(url)
# Drop fragments and common analytics parameters. Keep other query parameters
# because they may define real page variants on the target site.
query = [(key, value) for key, value in parse_qsl(parts.query, keep_blank_values=True)
if not key.lower().startswith("utm_") and key.lower() not in {"gclid", "fbclid"}]
query.sort()
path = re.sub(r"/{2,}", "/", parts.path or "/")
return urlunsplit((parts.scheme.lower(), parts.netloc.lower(), path, urlencode(query, doseq=True), ""))
def collect_urls(start_url):
pending = [start_url]
seen_sitemaps = set()
pages = set()
while pending:
sitemap = pending.pop()
if sitemap in seen_sitemaps:
continue
seen_sitemaps.add(sitemap)
kind, locations = fetch_sitemap(sitemap)
if kind == "sitemapindex":
pending.extend(urljoin(sitemap, item) for item in locations)
elif kind == "urlset":
for item in locations:
normalized = canonicalize(urljoin(sitemap, item))
host = (urlsplit(normalized).hostname or "").lower()
if host in ALLOWED_HOSTS:
pages.add(normalized)
if len(pages) > MAX_PAGES:
raise RuntimeError(f"URL count exceeds MAX_PAGES={MAX_PAGES}")
else:
raise ValueError(f"Unsupported sitemap root element: {kind}")
return sorted(pages)
def file_key(url):
parts = urlsplit(url)
path = parts.path.strip("/") or "home"
safe = re.sub(r"[^A-Za-z0-9._-]+", "_", path)[:100]
digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:12]
return f"{safe}-{digest}"
def main():
run_time = datetime.now(timezone.utc)
run_id = run_time.strftime("%Y%m%dT%H%M%SZ")
run_dir = OUTPUT_DIR / run_id
run_dir.mkdir(parents=True, exist_ok=True)
manifest_path = run_dir / "manifest.jsonl"
urls = collect_urls(SITEMAP_URL)
print(f"Found {len(urls)} unique sitemap URLs")
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
context = browser.new_context(
viewport={"width": 1440, "height": 1000},
device_scale_factor=1,
locale="en-US",
color_scheme="light",
)
page = context.new_page()
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
with manifest_path.open("w", encoding="utf-8") as manifest:
for index, url in enumerate(urls, start=1):
record = {
"source_url": url,
"captured_at": datetime.now(timezone.utc).isoformat(),
"viewport": {"width": 1440, "height": 1000},
"browser": "chromium",
}
try:
response = page.goto(url, wait_until="domcontentloaded")
record["final_url"] = page.url
record["status"] = response.status if response else None
if READY_SELECTOR:
page.locator(READY_SELECTOR).wait_for(state="visible", timeout=NAVIGATION_TIMEOUT_MS)
else:
page.wait_for_load_state("networkidle", timeout=10000)
# Scroll through the document to encourage lazy-loaded content
# to appear before taking the full-page screenshot.
page.evaluate("""async () => {
const step = Math.max(400, window.innerHeight);
for (let y = 0; y < document.body.scrollHeight; y += step) {
window.scrollTo(0, y);
await new Promise(resolve => setTimeout(resolve, 150));
}
window.scrollTo(0, 0);
}""")
page.screenshot(path=str(run_dir / f"{file_key(url)}.png"), full_page=True, animations="disabled")
record["result"] = "captured"
except Exception as error:
record["result"] = "error"
record["error"] = str(error)
try:
record["final_url"] = page.url
except Exception:
pass
manifest.write(json.dumps(record, ensure_ascii=False) + "\\n")
manifest.flush()
print(f"[{index}/{len(urls)}] {record['result']} {url}")
time.sleep(DELAY_SECONDS)
context.close()
browser.close()
if __name__ == "__main__":
main()
Set SITEMAP_URL and ALLOWED_HOSTS to the target site. For a small local run:
export SITEMAP_URL="https://example.com/sitemap.xml"
export ALLOWED_HOSTS="example.com,www.example.com"
export MAX_PAGES=100
python capture_sitemap.py
The script records a row for failures as well as successful captures, so a partial run remains reviewable. Its one-page-at-a-time browser loop favors a predictable request rate over throughput. Tune the delay and page limit based on target authorization, site capacity, and the number of URLs.
3. Schedule it with GitHub Actions
Commit the script and dependency file to a repository, then create .github/workflows/sitemap-screenshots.yml:
name: Scheduled sitemap screenshots
on:
workflow_dispatch:
schedule:
- cron: "17 3 * * *"
jobs:
capture:
runs-on: ubuntu-latest
timeout-minutes: 360
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: python -m pip install playwright requests
- run: python -m playwright install --with-deps chromium
- name: Capture sitemap pages
env:
SITEMAP_URL: ${{ vars.SITEMAP_URL }}
ALLOWED_HOSTS: ${{ vars.ALLOWED_HOSTS }}
MAX_PAGES: "50000"
DELAY_SECONDS: "1"
run: python capture_sitemap.py
- uses: actions/upload-artifact@v4
if: always()
with:
name: sitemap-screenshots-${{ github.run_id }}
path: captures/
retention-days: 14
This example runs daily at 03:17 UTC, deliberately away from the start of the hour. GitHub scheduled workflows use UTC and run against the latest commit on the default branch. The shortest schedule interval is five minutes, but schedules are not exact: GitHub says runs can be delayed during periods of high load and can occasionally be dropped. If a missed or late run has operational consequences, choose an execution platform and service guarantees that fit that requirement. GitHub Actions schedule documentation.
Use system cron instead
On a server you control, run the capture inside a virtual environment and keep logs:
17 3 * * * cd /opt/sitemap-shot && .venv/bin/python capture_sitemap.py >> /var/log/sitemap-shot.log 2>&1
Set the host timezone deliberately; this cron expression follows the machine’s local timezone. Add monitoring for nonzero exit status and stale output. A scheduled command alone does not retry failed pages or alert you that a run was missed.
4. Make captures comparable and useful
- Fix the rendering environment. Keep browser version, operating system, viewport, device scale factor, locale, timezone, color scheme, and fonts consistent between runs.
- Choose readiness per site. Prefer a meaningful selector such as a page’s main content container when possible. A universal fixed sleep is either wasteful or too short.
networkidlecan be inappropriate for pages with long polling, analytics, or continuously active connections. - Keep screenshots and manifests together. Record source and final URLs, timestamp, status, viewport, browser version, and error. The example records most of these; add fields for your own review process.
- Decide retention. Date-based folders make runs easy to compare, but image archives grow continuously. Keep a defined number of runs or move older artifacts to durable object storage.
- Review unstable content. Clocks, rotating banners, ads, personalized content, and animations can cause noisy diffs. Mask or disable such regions only when doing so reflects what you want to monitor.
Playwright Test can compare screenshots against baselines. Its documentation warns that rendering can vary with operating system, browser version, hardware, power conditions, and headless mode. Generate and compare baselines in the same environment, and review diffs before accepting a new baseline. Playwright visual comparisons.
A sitemap index entry’s lastmod describes when that sitemap file changed; it is not necessarily the update time for each page in that sitemap. Do not treat it as a page-level change signal unless you have verified the site’s data model. Sitemaps protocol.
5. Control load, scale, and reliability
Browser captures load page scripts, stylesheets, fonts, images, and other resources. Full-page captures may also need scrolling to reveal lazy-loaded content. A large sitemap can therefore create meaningful load on the target site and take a long time to process.
- Start with a small
MAX_PAGESand a conservative delay. Measure your own run time and adjust with the site owner’s authorization. - Limit concurrency. More browsers can shorten runs but also multiply simultaneous page loads and memory use. Add concurrency only after observing target impact and runner capacity.
- Split very large inventories into batches, such as one child sitemap per job. Preserve a manifest per batch so retries do not require recapturing successful pages.
- Retry transient network errors with a small bounded retry count and backoff. Do not retry permanent HTTP errors indefinitely.
- Record run start, finish, URL count, success count, failure count, and last successful run. Alert on missing or unexpectedly incomplete runs.
- Check storage and artifact retention. Full-page PNGs can consume substantial space; choose JPEG or WebP only if your comparison workflow accepts their compression behavior.
There is no universal safe pages-per-minute rate. A scanner project documents its own default of 60 pages per minute and allows a lower rate, but that is a project-specific setting, not a general recommendation. Follow site policies, honor applicable robots directives, and use a rate agreed with the site owner. Scanner project documentation.
6. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| No URLs found | Wrong sitemap URL, sitemap fetch blocked, or XML root not a sitemap index or URL set. | Open the sitemap URL directly, check the response and XML root, and confirm the URL is accessible from the runner. |
| Some child sitemaps fail | Index entries point to unavailable, disallowed, or malformed files. | Log the child sitemap URL and HTTP error; check access and retry transient failures separately. |
| URLs are missing from captures | Host allowlist excludes a host, URL normalization removed meaningful distinctions, or page cap was reached. | Review the allowlist and normalization rules; retain query parameters that identify real pages and raise the cap only when appropriate. |
| Navigation times out | Slow server, blocked resources, or page never reaches the chosen readiness state. | Increase timeout selectively, use a page-specific selector, or capture after domcontentloaded if later network activity never settles. |
| Screenshot is blank or incomplete | Capture happened before app rendering, page requires interaction, or lazy content was not triggered. | Wait for a meaningful selector, add the required authorized interaction, and verify scrolling triggers the content. |
| Images or fonts differ between runs | Unstable environment, cache state, or external assets changed. | Pin the browser/runtime environment where possible and record its version; distinguish external content changes from layout regressions. |
| GitHub run did not start on time | Scheduled workflows can be delayed or dropped under load. | Avoid the top of the hour, monitor run freshness, and use infrastructure with appropriate service guarantees for strict timing. |
| Runner runs out of memory or time | Too many large pages, excessive concurrency, or oversized full-page images. | Reduce concurrency, split sitemap batches, cap page count per job, and store results incrementally. |
| Artifact upload is empty | Output path differs or the script failed before creating the run directory. | Check the working directory and logs; upload the actual output path and retain logs even when capture fails. |
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-request API can capture URLs from your sitemap without maintaining a browser runtime. First extract and deduplicate the sitemap URLs as above, then call the API for each URL, respecting your desired rate and storing each response with a manifest. The ScreenshotNeo API documentation covers the request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents. - The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
For scheduled sitemap runs, keep the schedule and URL inventory logic in your workflow, then use the API for each capture. Sign up for 1,000 free screenshots a month, with no card required.
8. Frequently asked questions
Does a sitemap contain every page on a website?
No. It lists URLs the site has chosen to include. Treat the capture set as sitemap-listed pages; add a separate authorized crawl or test suite if you need route discovery.
Should I capture viewport screenshots or full pages?
Use full-page screenshots to review the entire scrollable document. Use viewport screenshots when the question concerns the initial visible layout or a particular device size.
Can I capture several viewports for each URL?
Yes. Repeat the capture for each viewport and include viewport dimensions in the filename or manifest so the outputs remain distinct and comparable.
Can screenshots prove every page works?
No. A screenshot records a rendered state. It does not prove that links, forms, menus, authentication flows, or other interactions work; test those separately.
Can I use sitemap lastmod to skip unchanged pages?
Only if the site’s sitemap data is trustworthy for that purpose. In a sitemap index, lastmod refers to the sitemap file, not necessarily each contained page.


