How to Ask an AI Agent to Screenshot Every Page in a Website Sitemap
Give an AI agent a defined sitemap scope, capture settings, and reporting format. This guide includes a reusable prompt and a runnable Playwright workflow.
To ask an AI agent to screenshot every page in a website sitemap, define the sitemap scope, how to handle URLs and redirects, the screenshot format and viewport, and what the agent must report for every URL. Then have it build a deduplicated manifest, visit each URL, save a stable-named screenshot, and reconcile the results against the manifest.
A sitemap is an inventory source, not proof that it contains every possible page or application state. “Every page” should mean every URL in the scope you define and successfully process—not every query variant, logged-in view, or route the sitemap never lists.
1. Define “every page” before you start
Fill in these decisions before giving the task to an agent. They prevent an ambiguous request from silently producing a partial or inconsistent set of screenshots.
- Site and allowed hosts: State the starting site and which hosts the agent may follow from a sitemap index. Do not let it follow arbitrary external sitemap references.
- Sitemap locations: Provide known sitemap URLs, or ask the agent to locate the site’s public sitemap and report which files it found. Sitemap indexes can refer to other sitemap files.
- URL policy: Say whether to preserve query strings, fragments, and locale paths. Fragments generally identify a position or state within a page rather than a separate server URL; decide whether those states are in scope.
- Redirects and canonical destinations: Decide whether to capture the URL as listed, the final destination after redirects, or both. Record the requested and final URL either way. Sitemaps generally describe preferred canonical URLs, not every alternate URL for the same content.
- Exclusions and access: List paths to skip. Ask the agent to inspect the applicable
robots.txtrules and stop or report if the requested scope is blocked or requires credentials it does not have. A robots.txt file applies to its host; its rules do not by themselves settle authorization for every kind of browser automation. - Capture mode: Choose a viewport screenshot for the visible screen, or a full-page screenshot for the whole scrollable document. These are different output goals.
- Consistent rendering: Pick one viewport, device scale, image format, browser state, and readiness condition for the run. A fixed delay may be appropriate for a specific page, but should not be assumed to suit every page.
- Output and accountability: Specify an output directory, filename policy, and a per-URL result record with requested URL, final URL, status, screenshot path, and error or exclusion reason.
Google recommends choosing the preferred URL when the same content is available at multiple URLs. It also cautions that submitting a sitemap is a hint and does not guarantee Google will download it or use it for crawling. That is another reason to report the exact manifest processed rather than claiming the sitemap proves complete site coverage. See Google’s sitemap guidance and robots.txt documentation.
2. Copyable prompt for an AI agent
Replace every bracketed value before running the task. Keep the scope and output rules in the prompt, even if the agent can infer some of them from context.
Inspect the public sitemap or sitemap index for [site URL] and build a manifest of the URLs in scope.
Scope
- Allowed hosts: [authorized hostnames]
- Sitemap URLs to inspect, if known: [sitemap URLs or “locate and report”]
- Follow sitemap index references only when they belong to an allowed host.
- Query-string policy: [preserve all / preserve these keys / ignore these keys]
- Locale policy: [include all locale paths / include these locales]
- Redirect policy: [capture listed URL / capture final destination / capture both]
- Exclusions: [paths or “none”]
- Authentication: [public pages only / describe credentials and permitted use]
Before browsing, inspect the applicable robots.txt rules. Stop and report if the requested scope is blocked or requires access you do not have. Do not expand the task to other hosts or routes without authorization.
Manifest and deduplication
- Follow in-scope sitemap index references and collect listed page URLs.
- Normalize URLs using the policies above and remove duplicates.
- Preserve meaningful query strings and locale paths as specified.
- Keep the source sitemap for each URL where practical.
- Do not infer that the sitemap represents every possible route or application state.
Capture
- Visit each manifest URL with Playwright.
- Capture mode: [viewport / full-page]
- Viewport: [width] x [height] CSS pixels; device scale factor: [value]
- Format: [PNG / JPEG / WebP]
- Output directory: [directory]
- Wait for [page-specific readiness condition]. Use an appropriate stable-render condition; do not apply one arbitrary delay to every page.
- Name files deterministically from a sanitized URL path plus a short collision-resistant suffix. Prevent collisions between URLs that differ by host, query, or normalization.
- Record both the requested URL and final URL after redirects.
Results
For every manifest entry, record requested URL, final URL, capture status, screenshot path (if any), and error or exclusion reason. Report counts for discovered, deduplicated, attempted, successful, redirected, failed, and excluded URLs. List every failure and exclusion.
Do not claim that every page on the site was captured. State that the result covers the defined sitemap manifest, and note that a sitemap may omit pages or list only preferred canonical URLs.
The prompt makes discovery and capture separate tasks: the sitemap supplies URLs to consider, and the browser must still navigate to each URL and capture it. Playwright’s full-page option means the full scrollable area of the current page; it does not visit every URL in a sitemap.
3. Run a local Playwright capture workflow
If the agent can use a terminal and a browser, ask it to implement or run a small script rather than relying on an undocumented sequence of manual actions. The following example is a runnable baseline for a known sitemap URL. It handles a sitemap index, filters to explicitly allowed hosts, deduplicates URLs, captures each URL, and writes a JSONL result record per entry.
It uses Playwright’s Node.js package. Install Node.js, then in a new working directory run:
npm init -y
npm install playwright
npx playwright install chromium
Save this as capture-sitemap.mjs. Replace the example sitemap and allowed hostname before running. The script defaults to viewport screenshots; use --full-page for full scrollable pages.
import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, writeFile, appendFile } from 'node:fs/promises';
import path from 'node:path';
const sitemapUrl = process.env.SITEMAP_URL ?? 'https://example.com/sitemap.xml';
const allowedHosts = new Set((process.env.ALLOWED_HOSTS ?? 'example.com')
.split(',').map(host => host.trim().toLowerCase()).filter(Boolean));
const outDir = process.env.OUT_DIR ?? 'shots';
const fullPage = process.argv.includes('--full-page');
const format = process.env.FORMAT ?? 'png'; // png, jpeg, or webp
const width = Number(process.env.WIDTH ?? 1365);
const height = Number(process.env.HEIGHT ?? 900);
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);
if (!['png', 'jpeg', 'webp'].includes(format)) throw new Error('FORMAT must be png, jpeg, or webp');
if (!Number.isInteger(width) || !Number.isInteger(height) || width <= 0 || height <= 0) {
throw new Error('WIDTH and HEIGHT must be positive integers');
}
await mkdir(outDir, { recursive: true });
const resultsPath = path.join(outDir, 'results.jsonl');
await writeFile(resultsPath, '');
function inScope(raw) {
try {
const u = new URL(raw);
return u.protocol === 'https:' || u.protocol === 'http:' ? allowedHosts.has(u.hostname.toLowerCase()) : false;
} catch { return false; }
}
function filenameFor(raw) {
const u = new URL(raw);
const readable = (u.hostname + u.pathname).replace(/[^a-zA-Z0-9._-]+/g, '-').replace(/-+/g, '-').slice(0, 100) || 'page';
const suffix = createHash('sha256').update(raw).digest('hex').slice(0, 10);
return `${readable}-${suffix}.${format === 'jpeg' ? 'jpg' : format}`;
}
async function fetchXml(url) {
const response = await fetch(url, { signal: AbortSignal.timeout(timeoutMs) });
if (!response.ok) throw new Error(`Sitemap HTTP ${response.status}: ${url}`);
return await response.text();
}
function sitemapEntries(xml, tag) {
const re = new RegExp(`<${tag}(?:\\s[^>]*)?>\\s*<loc>([^<]+)</loc>\\s*</${tag}>`, 'gi');
return [...xml.matchAll(re)].map(m => m[1].trim().replaceAll('&', '&'));
}
async function collect(url, visited = new Set()) {
const normalized = new URL(url).href;
if (visited.has(normalized)) return [];
visited.add(normalized);
const xml = await fetchXml(normalized);
const childSitemaps = sitemapEntries(xml, 'sitemap');
if (childSitemaps.length) {
const urls = [];
for (const child of childSitemaps) {
if (!inScope(child)) continue;
urls.push(...await collect(child, visited));
}
return urls;
}
return sitemapEntries(xml, 'url').filter(inScope);
}
const discovered = await collect(sitemapUrl);
const manifest = [...new Set(discovered.map(raw => new URL(raw).href))];
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width, height } });
let successes = 0, failures = 0, redirects = 0;
try {
for (const requestedUrl of manifest) {
const page = await context.newPage();
let record = { requestedUrl, finalUrl: null, status: null, screenshotPath: null, error: null };
try {
const response = await page.goto(requestedUrl, { waitUntil: 'networkidle', timeout: timeoutMs });
record.finalUrl = page.url();
record.status = response?.status() ?? null;
if (response && response.status() >= 400) throw new Error(`Page HTTP ${response.status()}`);
const target = path.join(outDir, filenameFor(requestedUrl));
await page.screenshot({ path: target, fullPage, type: format });
record.screenshotPath = target;
successes++;
if (record.finalUrl !== requestedUrl) redirects++;
} catch (error) {
record.finalUrl = page.url();
record.error = error instanceof Error ? error.message : String(error);
failures++;
} finally {
await appendFile(resultsPath, JSON.stringify(record) + '\n');
await page.close();
}
}
} finally {
await context.close();
await browser.close();
}
console.log(JSON.stringify({ discovered: discovered.length, deduplicated: manifest.length,
attempted: manifest.length, successful: successes, redirected: redirects, failed: failures,
excluded: 0, resultsPath }, null, 2));
Run it with an explicit scope and output settings:
SITEMAP_URL=https://example.com/sitemap.xml \
ALLOWED_HOSTS=example.com,www.example.com \
OUT_DIR=shots WIDTH=1365 HEIGHT=900 FORMAT=png \
node capture-sitemap.mjs
For full-page screenshots, append --full-page. The script uses networkidle as a baseline, but that can be a poor readiness signal on pages with long polling, analytics, or continuously active connections. For those sites, adapt the wait condition—for example, wait for a page-specific selector or a documented load state—and record that policy. Playwright documents viewport, element, and full-page screenshots, custom filenames, and PNG/JPEG/WebP output in its screenshot guide. Its quick start shows the basic browser interaction loop.
Important limits of the sample
- The small XML parser is intended for ordinary sitemap XML. For production runs, use a proper XML parser, especially if the files use namespaces, unusual formatting, or entity encoding.
- The sample follows same-allowlist sitemap index references. Confirm that each allowed host and sitemap is in scope before running.
- It does not inspect robots.txt, authenticate, or implement retries. Do those as explicit, site-specific steps; do not silently retry indefinitely.
- The URL deduplication uses normalized URL strings, so URLs with different query strings remain distinct. Change that only if your stated policy says those query variations are equivalent.
- A status of
nullcan occur when navigation has no main document response, such as some client-side transitions. Treat it as something to inspect rather than proof of success. - Full-page screenshots of very long documents can be large. If image review becomes unwieldy, define a section or tiled capture policy in advance and record it.
4. Make the run auditable
Compare the result records with the final manifest. A useful completion report includes these counts and a list of exceptions:
| Count | Meaning |
|---|---|
| Discovered | URL entries read from the inspected sitemap files before deduplication and scope filtering. |
| Deduplicated | Unique in-scope URLs after applying the declared normalization policy. |
| Attempted | Manifest entries for which the browser navigation and capture process began. |
| Successful | Entries with a saved screenshot and no reported capture error. |
| Redirected | Entries whose final URL differs from the requested URL. Report whether the screenshot represents the source or destination according to the declared policy. |
| Failed | Entries where navigation, readiness, or screenshot saving failed. |
| Excluded | Entries intentionally skipped by the stated scope, with a reason for each. |
Also state the sitemap files inspected, capture mode, viewport, format, readiness rule, and the timestamp or run identifier. Keep one result row for every manifest entry, including failures. This makes it possible to distinguish a complete run of a defined inventory from a claim about the whole website.
5. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Only a fraction of expected pages appear | The sitemap index was not followed, a child sitemap was outside the host allowlist, or the sitemap intentionally lists only canonical URLs. | Report all inspected sitemap files, check the index references and host policy, and define whether another authorized inventory source is in scope. |
| Duplicate screenshots or filename collisions | Different URLs were reduced to the same sanitized path, or the filename omitted host/query distinctions. | Include a collision-resistant hash of the complete normalized URL, as the sample does. Keep meaningful query strings distinct under the declared policy. |
| Pages time out waiting for network idle | The page keeps network connections active or has long-running requests. | Use a page-specific readiness condition or a suitable load event, and record the choice. A longer timeout alone may not solve a condition that never occurs. |
| Screenshot is blank or incomplete | The page had not rendered meaningful content, lazy content had not loaded, a script failed, or the selected readiness condition was too early. | Inspect the page and console/network errors, wait for a meaningful selector or content condition, and decide whether scrolling is needed to trigger lazy loading. |
| HTTP error or navigation failure | The URL is unavailable, blocked, redirected unexpectedly, or the server returned an error. | Record the response status and final URL. Do not treat an error page as a successful capture unless error pages are explicitly part of the requested deliverable. |
| Sitemap fetch fails | Wrong sitemap URL, inaccessible file, invalid XML, or a child sitemap outside the authorized scope. | Report the failed sitemap and reason. Verify the URL and XML, and do not expand to another host without authorization. |
| Many pages look different between runs | Responsive breakpoints, dynamic content, consent state, personalization, locale, time, or browser settings changed. | Hold viewport, device scale, browser context, locale, and other relevant state constant; document any state that cannot be controlled. |
6. Performance, reliability, and cost
A sitemap batch makes one browser navigation per URL, plus sitemap retrieval. Total work grows with the number of distinct URLs and the time each page needs to reach the chosen readiness condition. Full-page captures may use more memory and produce larger files than viewport captures, especially for tall pages.
- Start with bounded concurrency. The sample processes one page at a time. If increasing concurrency, do so gradually and respect the site’s capacity and scope; no universal concurrency setting is appropriate for every site.
- Use explicit timeouts and finite retries. A timeout makes stalled pages visible. If you add retries, bound the count, record each attempt, and avoid repeated traffic against a failing site.
- Keep the manifest durable. Save it before capture, write results incrementally, and make reruns able to identify completed entries. This reduces lost work if a long job stops.
- Estimate storage from a sample. Capture a small representative set first, note the chosen mode and format, then estimate storage for the manifest. Actual image sizes depend on page content and dimensions.
- Control state deliberately. A fresh browser context gives a consistent starting point, but sites may still vary by time, location, experiment, or server-side personalization.
- Budget browser resources. Each browser page consumes resources. Close pages after capture, and avoid opening the whole sitemap at once.
For local Playwright runs, the direct costs are your compute, storage, and network use; the exact amounts depend on the machine, site, image output, and run. For hosted capture, check the provider’s billing rules for failed pages, caching, and retries before choosing it.
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. For one URL, call the API with an access key and target URL; see the API documentation for capture parameters. This one-call example saves the response as a WebP file:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
For an AI agent processing a sitemap, it can call the endpoint once for each URL in the manifest and save each response under the deterministic filename recorded in its result manifest. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing outcome. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
8. FAQ
Does a sitemap list literally every page?
No. It is the inventory defined by the sitemap publisher and may list preferred canonical URLs while omitting query states, logged-in pages, or other routes. Report the sitemap scope you actually processed.
Should the agent capture full-page screenshots?
Use full-page mode when reviewers need the entire scrollable document in one image. Use viewport mode when the review concerns the initial visible presentation or a specific screen size.
Should redirected URLs count as successful?
That depends on the requested deliverable. Record both the requested and final URLs, then apply the redirect policy you chose before the run.
Can this workflow cover authenticated pages?
Only if access and credentials are explicitly authorized and supplied through an appropriate secure method. A public sitemap alone does not provide access to authenticated content.
Primary references: Playwright screenshots, Playwright quick start, Google Search Central sitemap guidance, and Google robots.txt documentation.


