How to Capture Bulk Website Screenshots as PDFs with Playwright
Use Playwright to capture a URL list as separate PDFs or full-page images, with reliable readiness checks, print settings, and controlled concurrency.
To capture many websites as PDFs with Playwright, make a URL list, open each URL in Chromium, wait for the content that matters, and call page.pdf() with a unique output path. A PDF uses print CSS by default. For a full-page image instead, call page.screenshot({ fullPage: true }); that produces a tall raster image, not a paginated PDF.
The example below saves one A4 PDF per URL, uses a deterministic filename, records failures without stopping the whole batch, and closes each page after capture. PDF generation is supported in Chromium; check that your installed Playwright version and browser runtime support the API you use. See the official Playwright Page PDF API and screenshot API.
1. Install Playwright and prepare a URL list
Use a current Node.js runtime and install Playwright in a project. Install its Chromium browser as well:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as capture-pdfs.mjs. Replace the sample URLs with the pages you are authorized to capture. The script writes files to a captures directory next to the script and continues after individual navigation or PDF errors.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
const here = path.dirname(fileURLToPath(import.meta.url));
const outputDir = path.join(here, 'captures');
const urls = [
'https://example.com/',
'https://playwright.dev/',
];
function outputName(url, index) {
const host = new URL(url).hostname.replace(/[^a-z0-9.-]/gi, '_');
return `${String(index + 1).padStart(3, '0')}-${host}.pdf`;
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const context = await browser.newContext();
const failures = [];
try {
for (const [index, url] of urls.entries()) {
const page = await context.newPage();
const destination = path.join(outputDir, outputName(url, index));
try {
const response = await page.goto(url, {
waitUntil: 'load',
timeout: 45_000,
});
if (!response) {
throw new Error('Navigation returned no main-resource response');
}
if (!response.ok()) {
throw new Error(`Main-resource HTTP status ${response.status()}`);
}
// Replace or supplement this with a condition meaningful to the target site.
await page.locator('body').waitFor({ state: 'visible', timeout: 10_000 });
await page.pdf({
path: destination,
format: 'A4',
printBackground: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
});
console.log(`Saved ${destination}`);
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
failures.push({ url, error: message });
console.error(`Failed ${url}: ${message}`);
} finally {
await page.close();
}
}
} finally {
await context.close();
await browser.close();
}
if (failures.length) {
const report = path.join(outputDir, 'failures.json');
await writeFile(report, JSON.stringify(failures, null, 2));
console.error(`${failures.length} URL(s) failed. See ${report}`);
process.exitCode = 1;
}
Run it with:
node capture-pdfs.mjs
The response check catches HTTP error statuses for the main document, but it does not guarantee that every subresource loaded successfully or that a single-page application finished rendering. Add site-specific checks for those requirements.
2. Choose a readiness condition
Navigation completion and page readiness are different. load waits for the page load event, but client-side rendering, API calls, fonts, animations, and lazy content can continue afterward. Choose a wait condition based on what the page must contain before capture.
- Known element: wait for a content selector, such as
main articleor a page-specific heading. - Known state: wait for a URL change, a particular text value, or an application-owned ready marker.
- Lazy-loaded content: scroll through the page before printing, then wait for images or other content to finish.
- Fixed delay: use only when the site has no observable readiness condition. A delay is less reliable and can waste time.
For example, replace the generic body wait with a selector that confirms meaningful content:
await page.locator('main article h1').waitFor({ state: 'visible', timeout: 15_000 });
For a site with an application-specific readiness signal, wait for it explicitly:
await page.waitForFunction(() => window.document.documentElement.dataset.ready === 'true', null, {
timeout: 15_000,
});
That example assumes the target page sets data-ready="true" on its root element. Use a real signal exposed by the site; do not assume every page provides one.
3. Configure PDF layout and styling
page.pdf() renders with print CSS by default. Print styles can hide navigation, change columns, or replace screen-oriented layouts. If you want the page’s screen styling, emulate screen media before generating the PDF:
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-style.pdf', format: 'A4', printBackground: true });
Use print media when the website provides an intentional print layout. Use screen media when you need the on-screen design. Check representative pages in the resulting files because a screen layout may paginate awkwardly on paper.
| Option | What it controls | Practical guidance |
|---|---|---|
format |
Standard paper size such as A4 or Letter | Use when output should follow a familiar paper size. |
width, height |
Custom page dimensions | Use instead of a standard format when the document needs specific dimensions. |
margin |
Top, right, bottom, and left page margins | Set explicitly for consistent output; CSS page rules may also affect layout. |
printBackground |
Whether background graphics and colors are printed | Enable it when backgrounds carry visual information. |
scale |
Scales the page content | Adjust carefully; reducing scale can make text hard to read. |
pageRanges |
Which PDF pages to include | Useful for extracting a known range from long output. |
preferCSSPageSize |
Whether CSS @page dimensions take precedence |
Enable when the site’s print stylesheet defines the intended paper size. |
landscape |
Paper orientation | Use when wide tables or layouts need landscape pages. |
When both PDF options and page CSS specify dimensions, the chosen settings can interact. In particular, consider @page rules and preferCSSPageSize together. The official PDF API reference documents the available options and their behavior.
Example: print backgrounds and preserve intended colors
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '10mm', right: '10mm', bottom: '10mm', left: '10mm' },
});
For exact print colors, the page’s CSS can use -webkit-print-color-adjust. This is a stylesheet rule, not a Playwright PDF option. If you control the page, add an appropriate print rule; otherwise, inspect whether its existing print styles produce the colors you need.
4. Capture full-page images instead of PDFs
A PDF is a paper-sized document with page breaks. A full-page screenshot is one raster image of the page’s scrollable area. Use page.screenshot() when the consumer needs an image, not a paginated document.
await page.screenshot({
path: 'full-page.png',
fullPage: true,
animations: 'disabled',
scale: 'css',
});
Screenshot options include output path, image type, quality for JPEG, full-page capture, clipping, scale, and animation behavior. The exact combinations depend on the selected image type; consult the screenshot API. Full-page images can become very large for long pages. For a viewport capture, omit fullPage or set it to false. To capture a specific element, use the element screenshot API on a locator.
Full-page mode captures the scrollable page area, but it does not guarantee that every lazy-loaded image or infinite-scroll item has been rendered. If complete content matters, scroll deliberately and wait for the specific content before taking the screenshot.
5. Process larger batches carefully
The sequential example is a good starting point for modest batches: it limits simultaneous pages and makes failures easier to trace. Playwright contexts can host multiple pages, and a context exposes its pages through context.pages(). The documentation does not prescribe a universal safe concurrency value. Site complexity, available memory, CPU, browser settings, and the runtime all affect a suitable limit.
If a sequential run takes too long, use a bounded worker pool rather than launching every URL at once. Each worker should have its own page, and the number of workers should be a configuration value you can tune after observing resource use in the target environment.
const concurrency = 3; // Start conservatively; tune for your environment.
let nextIndex = 0;
async function worker() {
while (true) {
const index = nextIndex++;
if (index >= urls.length) return;
const page = await context.newPage();
try {
const response = await page.goto(urls[index], { waitUntil: 'load', timeout: 45_000 });
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 });
await page.pdf({ path: outputPathFor(urls[index], index), format: 'A4', printBackground: true });
} finally {
await page.close();
}
}
}
await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, () => worker()));
This is a worker pattern to integrate into the earlier script; define outputPathFor using the same filename logic and add per-URL error collection if you want a failed page not to reject the whole Promise.all. Avoid duplicate output paths. If you need better failure isolation, have each worker catch and record errors around each URL, as the sequential example does.
Repeatability checklist
- Pin the Playwright package and install the matching browser in the environment used for capture.
- Keep the host OS, browser mode, viewport, locale, timezone, and other relevant settings consistent.
- Use explicit readiness conditions instead of relying on arbitrary short delays.
- Use deterministic names and record URL-to-file results, including failures.
- Review representative PDFs after changing browser versions or runtime settings.
Rendering can vary with host operating system, browser version and settings, hardware, power source, and headless mode. A controlled environment improves repeatability, but review output when visual consistency matters. The Playwright browser documentation describes browser installation and supported browser usage.
6. Use the Playwright CLI for one-off captures
For an occasional capture, the Playwright CLI includes screenshot and PDF commands, including a full-page screenshot option. A script is a better fit when you need a repeatable URL-driven batch, per-site readiness checks, custom output names, or failure reporting. See the Playwright CLI documentation for the commands and options supported by your installed version.
7. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
page.pdf is not a function or PDF generation fails |
The selected browser or runtime does not support the PDF API, or installed components are mismatched. | Use Chromium for PDF generation and install the browser version matching the Playwright package. Check the versioned Page API. |
| PDF is missing styles or looks unlike the browser | PDF uses print media by default, and the site may have print-specific CSS. | Choose print media intentionally, or call page.emulateMedia({ media: 'screen' }) before page.pdf(). |
| Background colors or graphics are absent | Background printing is disabled, or the site changes styling for print. | Set printBackground: true and inspect print CSS, including color adjustment rules if you control the site. |
| PDF content is blank or incomplete | Capture began before client-side rendering, or a main-resource response failed. | Check the navigation response and wait for a page-specific content selector or application readiness signal. |
| Images are missing | Images may be lazy-loaded, blocked, or not finished loading when capture starts. | Scroll to trigger lazy loading, wait for relevant images, and investigate failed image requests. Do not treat network-idle alone as proof every image is ready. |
| Navigation times out | The site is slow, continues network activity, or does not reach the selected lifecycle event. | Set a suitable timeout, choose a lifecycle event that fits the page, then wait on a meaningful content condition. Record the failure and retry selectively if appropriate. |
| Some URLs overwrite other PDFs | Output filenames collide, often because names use only the hostname. | Include the input index or a stable unique identifier in every output name. |
| Batch becomes slow or the process runs out of memory | Too many pages are open at once, or long pages produce large output. | Close pages after each capture, process sequentially or use a bounded worker pool, and reduce concurrency based on observed resource use. |
| Results differ between runs | Browser version, host environment, timing, dynamic content, or animation state changed. | Control the runtime and capture settings, wait on explicit readiness conditions, and disable or wait out animations when appropriate. |
8. Performance, reliability, and cost
Playwright is self-managed: you run the browser and pay for the compute, storage, and operational work associated with that environment. Total run time depends on the number of URLs, page load behavior, PDF complexity, and concurrency. There is no universal throughput figure or safe worker count. Start sequentially, measure on representative URLs, then increase bounded concurrency while watching resource use and failure rates.
For reliability, keep a per-URL result record, distinguish navigation failures from output failures, use timeouts, close pages in finally blocks, and rerun only failed items when possible. A successful PDF call means a file was generated; inspect output when correctness or visual fidelity matters. Dynamic pages can change between runs even when the capture code is unchanged.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a screenshot or PDF. Its clean-capture flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
For a PDF, call the API with format=pdf:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
timeout=90,
)
r.raise_for_status()
with open("page.pdf", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('page.pdf', Buffer.from(await res.arrayBuffer())));
See the ScreenshotNeo API documentation for request parameters and configuration. ScreenshotNeo supports bulk capture of up to 100 URLs per call, and its parameter names also work with those used by other screenshot APIs to make switching easier. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Create a free account and get 1,000 screenshots a month with no card.
FAQ
Can one Playwright page produce a PDF and an image?
Yes. After navigating and waiting for readiness, call page.pdf() and page.screenshot() as separate operations with different output paths. Choose each output’s layout and format independently.
Can I save one PDF per URL?
Yes. Create a page for each URL and give every output a unique path. The batch script above uses an index and hostname to avoid common filename collisions.
Does page.pdf() capture only the visible viewport?
No. It generates a paginated document according to print layout and PDF settings. For a single image of the full scrollable page, use page.screenshot({ fullPage: true }).
Can I capture authenticated pages?
Yes, when you provide the appropriate authentication state or credentials to the browser context and have permission to access the pages. Keep credentials out of source control and avoid sharing generated files that contain private content.


