How to Capture Screenshots of Every Page on a Website
Build a reliable URL inventory, capture each page with Playwright, and track failures so you know exactly what your website screenshot run covered.
To capture screenshots of every page on a website, first create and validate a list of the URLs you mean to include, then visit each URL with browser automation and save one screenshot per page. Playwright can capture either the visible viewport or the entire scrollable page with fullPage: true. That option captures one visited page; it does not discover the site’s URLs for you.
A reliable batch job therefore has two parts: URL discovery and page capture. Keep a manifest of intended URLs and outcomes so you can distinguish a complete run from one with skipped or failed pages.
1. Define what “every page” means
Before collecting URLs, decide which page states belong in the run. A site’s URL space may include public canonical pages, localized routes, query-string variants, authenticated pages, search results, and parameterized pages. Some patterns, such as calendars or faceted filters, can generate effectively unlimited URLs.
- Set the allowed hostnames and path prefixes.
- Decide whether query strings and trailing-slash variants represent separate pages.
- Decide whether login is in scope and how credentials will be provided.
- Exclude URL patterns that can create unbounded variants.
- Choose whether you need the initial viewport or the full scrollable page.
“Every page” is only meaningful relative to this scope. A screenshot run cannot prove that it found every possible route unless the URL inventory is complete for the site and purpose.
2. Build and validate the URL inventory
Start with the sitemap
Look for a sitemap at a known location such as /sitemap.xml, or inspect the site’s robots.txt for sitemap declarations. A sitemap index can point to multiple sitemap files. The sitemap protocol allows up to 50,000 URLs and 50 MB per sitemap file; larger inventories can be split across files. See the Sitemaps protocol.
Treat sitemap entries as candidate URLs, not proof of exhaustive coverage. Google describes sitemaps as a discovery aid and says they do not guarantee every listed item will be crawled and indexed. For a screenshot task, the practical lesson is to check the sitemap against your intended scope and other route sources. See Google’s sitemap overview.
Check for omissions
Depending on the site, supplement the sitemap with an internal-link crawl, a route list from the application, or a URL list supplied by the site owner. These sources can reveal unlinked pages or routes not represented in the sitemap. Do not follow every discovered link without scope controls: generated filters, pagination, and calendar dates can cause a crawler to expand indefinitely.
Normalize and constrain URLs
Before capture, remove duplicates and reject URLs outside the allowed host and path. Decide how to handle fragments, query parameters, redirects, and trailing slashes. Keep the validated inventory as a durable input file so a later run can use the same scope.
3. Capture every URL with Playwright
The example below uses Node.js and Playwright. It reads one URL per line from urls.txt, visits each URL, saves a full-page PNG, and writes a JSON Lines manifest with the requested URL, final URL, timestamp, and success or failure. It uses a fixed viewport and a bounded navigation timeout. Install Playwright and its Chromium browser using the official installation instructions; the Page screenshot API documents capture options.
// capture.mjs
import fs from 'node:fs/promises';
import path from 'node:path';
import { chromium } from 'playwright';
const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const manifestFile = path.join(outputDir, 'manifest.jsonl');
const allowedHosts = new Set((process.env.ALLOWED_HOSTS ?? '')
.split(',').map(s => s.trim().toLowerCase()).filter(Boolean));
const urls = (await fs.readFile(inputFile, 'utf8'))
.split(/\r?\n/).map(s => s.trim()).filter(s => s && !s.startsWith('#'));
await fs.mkdir(outputDir, { recursive: true });
await fs.writeFile(manifestFile, '');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1
});
const page = await context.newPage();
function filenameFor(url, index) {
const parsed = new URL(url);
const slug = (parsed.pathname.replace(/[^a-z0-9]+/gi, '-').replace(/^-|-$/g, '') || 'home')
.slice(0, 70);
return `${String(index + 1).padStart(5, '0')}-${slug}.png`;
}
async function record(entry) {
await fs.appendFile(manifestFile, `${JSON.stringify(entry)}\n`);
}
try {
for (let i = 0; i < urls.length; i++) {
const requestedUrl = urls[i];
const startedAt = new Date().toISOString();
let finalUrl = null;
try {
const parsed = new URL(requestedUrl);
if (!['http:', 'https:'].includes(parsed.protocol)) throw new Error('Only HTTP(S) URLs are allowed');
if (allowedHosts.size && !allowedHosts.has(parsed.hostname.toLowerCase())) {
throw new Error(`Host is outside ALLOWED_HOSTS: ${parsed.hostname}`);
}
const response = await page.goto(requestedUrl, {
waitUntil: 'domcontentloaded', timeout: 30000
});
finalUrl = page.url();
// Allow a brief paint interval for client-rendered content. Adjust for the site.
await page.waitForTimeout(500);
const outputPath = path.join(outputDir, filenameFor(requestedUrl, i));
await page.screenshot({ path: outputPath, fullPage: true, type: 'png' });
await record({ requestedUrl, finalUrl, status: response?.status() ?? null,
outcome: 'captured', file: outputPath, startedAt, finishedAt: new Date().toISOString() });
} catch (error) {
await record({ requestedUrl, finalUrl, outcome: 'failed', error: String(error),
startedAt, finishedAt: new Date().toISOString() });
}
}
} finally {
await browser.close();
}
Save the script as capture.mjs and create urls.txt with one absolute URL per line. Run it with:
npm install playwright
npx playwright install chromium
ALLOWED_HOSTS=example.com,www.example.com node capture.mjs urls.txt screenshots
On Windows PowerShell, set the host variable with $env:ALLOWED_HOSTS="example.com,www.example.com" before running the Node command. The host check is a useful guard when the URL list is generated or supplied dynamically. For a trusted, prevalidated list, the variable can be left empty.
4. Choose capture settings consistently
| Setting | Choice and effect |
|---|---|
| Viewport or full page | fullPage: false captures the current viewport; true captures the full scrollable page as a tall image. Full-page capture applies to one URL at a time. |
| Viewport dimensions | Set width and height before navigation for repeatable desktop or mobile-sized captures. The viewport can change responsive layout and content. |
| Device scale | deviceScaleFactor controls device-pixel scaling relative to CSS pixels. Keep it fixed when comparing images; a higher scale produces larger image dimensions and files. |
| Image format | PNG is lossless and useful for detailed review or pixel comparison. Playwright also supports JPEG and WebP screenshot output; quality applies to supported lossy formats. |
| Path and naming | Set a deterministic output path. Include an index or stable ID because different URLs can map to the same path-derived slug. |
| Masking | Playwright can mask matching locators in a screenshot, useful when variable page regions would otherwise make visual comparisons noisy. |
| Transparency/background | Background behavior depends on image format and screenshot options. Use a format and background configuration appropriate to the output you need. |
For comparison runs, keep browser version, viewport, device scale, color scheme, locale, authentication state, and readiness rule stable too. These can change what the page renders even when the URL is unchanged.
5. Handle dynamic pages and lazy-loaded content
A navigation event does not guarantee that every application has finished rendering. The example waits for DOM content and a short paint interval, but that is only a starting point. For a site you control, wait for a meaningful selector that signals the page is ready, such as the main content container. A fixed delay can be wasteful on fast pages and too short on slow ones.
Full-page capture records the scrollable page, but lazy-loaded images or content that appears only after scrolling may not load automatically. If those matter, scroll through the page in controlled increments, wait for content to settle, then capture; site-specific infinite-scroll pages need an explicit stopping rule. Avoid waiting for network idle blindly on pages with long-lived polling or analytics requests.
For authenticated routes, use a deliberate browser context and an authorized login state. Keep cookies and tokens out of source code and manifests. If the site blocks automation or requires an interactive challenge, follow its access rules rather than trying to bypass them.
6. Review the manifest and coverage
After the run, compare the manifest with the validated URL inventory. Count captured, failed, and intentionally skipped entries. Inspect non-2xx responses, redirects, blank images, challenge pages, and screenshots that are unexpectedly short or tall. A successful screenshot write only proves an image was saved; it does not prove that the intended content loaded.
Keep the original URL list, capture settings, timestamp, browser version, and manifest together. This makes reruns and audits easier and prevents a partial run from being mistaken for full coverage.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Some pages are missing | The sitemap or supplied list omitted routes; discovery and capture are separate. | Compare the inventory with internal links or the application’s route list, then rerun missing URLs. |
| Navigation timeout | Slow server, redirect loop, or wait condition that never settles. | Inspect the URL and redirect behavior, use a bounded timeout, and choose a readiness condition suited to the page. |
| Screenshot is blank or incomplete | Client rendering has not completed, content needs interaction, or the URL returned an error/challenge. | Check the response status and final URL, wait for a content selector, and inspect the page before changing delays. |
| Images are missing in a full-page capture | Lazy loading or scroll-triggered content has not been activated. | Scroll through the page before capture and wait for the relevant images or content to load. |
| Images differ between runs | Viewport, device scale, dynamic content, fonts, animation, or timing changed. | Fix capture settings and state; mask volatile regions when appropriate and wait for stable content. |
| Output files overwrite each other | Filename derived only from a path that is shared or normalized identically. | Include a sequence number or stable URL ID, as the sample does. |
| Access denied or challenge page | The site applies access controls, bot checks, or authentication. | Use authorized access and credentials where permitted; record the outcome and respect the site’s rules. |
| Browser launch fails | Playwright’s browser binary is not installed or runtime dependencies are missing. | Follow Playwright’s installation instructions and install the required browser for the environment. |
8. Performance, reliability, and cost
Browser rendering is usually the expensive part of a batch: every URL requires navigation, page execution, and image output. A single-page-at-a-time loop is simple and gives clear failure isolation. For larger inventories, controlled concurrency can reduce elapsed time, but too many simultaneous pages may exhaust memory or trigger site rate limits. Start with low concurrency, record failures, and retry transient errors selectively rather than restarting the entire inventory.
Full-page images can be very tall and large, especially at high device scale. Use viewport captures when only the first screen matters; choose JPEG or WebP when smaller lossy files are acceptable. Store output incrementally, keep a manifest, and ensure the destination has enough disk space. Run captures on a stable browser/runtime version when results need to be comparable.
With a local Playwright setup, account for the compute time, storage, and engineering effort to maintain browser dependencies, URL discovery, retries, and reporting. The research sources do not establish a universal capture speed or cost: these depend on the site, page weight, image dimensions, and execution environment.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL as PNG, JPEG, WebP, or PDF; its parameters include full-page capture for scrollable content. It does not discover every URL for you, so provide the validated URL inventory and call it once per page or use the bulk capture option for up to 100 URLs per call. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan.
Sign up free for 1,000 screenshots a month, no card required.
FAQ
Does a full-page screenshot find every route?
No. It captures the scrollable content of one URL. URL discovery and coverage validation are separate steps.
Can a sitemap guarantee that I have every page?
No. It is a useful candidate list, but check it against the routes and scope you actually need.
Should I capture full-page or viewport images?
Use full-page images for whole-document review or archives; use viewport images when the first screen or a consistent device view is the goal.
Can I compare screenshots pixel by pixel?
Yes, but only treat differences as meaningful when the viewport, scale, page state, and rendering conditions are controlled.


