How to Archive a Webpage as a Screenshot and PDF Automatically
Capture a webpage as a screenshot and PDF with Chrome Headless, Playwright, or Puppeteer. Learn how to wait for dynamic content and keep both files organized.
To archive a webpage automatically, open it in a browser automation tool, wait for the page state you need, then save a screenshot and a PDF. For a one-off capture, Chrome Headless can do both from the command line. Use Playwright or Puppeteer when you need full-page or element screenshots, per-page readiness checks, repeatable filenames, or a multi-URL job.
Save both files with a consistent name and record the source URL and capture time alongside them. A screenshot records a visual rendering, and a PDF records a print-oriented document; neither guarantees a permanent record of every interactive feature or resource on the live site.
1. Choose a capture method
| Method | Useful when | What to account for |
|---|---|---|
| Chrome Headless CLI | You need a straightforward URL-to-file command. | The command can wait up to a timeout, but the page may still be loading when it captures. |
| Playwright | You need viewport, full-page, or element screenshots, or a programmable workflow. | Choose an explicit readiness condition and output naming scheme. |
| Puppeteer | You want a programmable Chromium-based workflow and control over PDF media styling. | PDF uses print CSS by default. Select screen media first if you want the screen presentation. |
The documentation establishes these capabilities, not a speed ranking. Choose based on the capture controls and workflow you need.
2. Quick capture with Chrome Headless
Chrome Headless provides command-line switches for screenshots and PDF output. Run these commands from the directory where you want the files saved, substituting the target URL.
chrome --headless --screenshot=page.png https://example.com
chrome --headless --print-to-pdf=page.pdf https://example.com
To set a maximum wait before capture, use --timeout with a value in milliseconds:
chrome --headless --timeout=10000 --screenshot=page.png https://example.com
The timeout is a maximum wait, not proof that scripts, images, or embedded content have finished. Chrome may capture while the page is still loading. See the Chrome Headless command-line reference for the documented switches and behavior.
3. Automate screenshots and PDFs with Playwright
Playwright supports viewport, full-page, and element screenshots. Its CLI also supports PDF capture. Install Playwright and its browser as described in the official installation guide, then use this Node.js script to save both outputs for a URL:
import { chromium } from 'playwright';
const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
try {
await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
await page.screenshot({ path: 'page-full.png', fullPage: true });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
await browser.close();
}
Run it with node archive.mjs https://example.com. If a site keeps network connections open, networkidle may not be an appropriate readiness condition. Use a page-specific selector or another documented navigation condition for your site instead. Playwright’s Page API documents navigation and page capture; its screenshots and PDF guide describes screenshot targets and CLI export.
For a single CLI PDF capture, Playwright documents:
playwright pdf https://example.com page.pdf
Viewport, full-page, and element output
A screenshot defaults to the visible viewport unless you request a larger extent. For a full scrollable page, use fullPage: true. To capture only a particular element, locate it and use its screenshot method:
const article = page.locator('article');
await article.screenshot({ path: 'article.png' });
Use a viewport screenshot when the initial screen is the record you need, full-page output when the entire page matters, and an element screenshot when you need a specific region. Long pages can produce large images; check the output before relying on it as a shareable record.
4. Automate screenshot and PDF capture with Puppeteer
Install Puppeteer following its official installation guide. This script captures a full-page screenshot and a PDF:
import puppeteer from 'puppeteer';
const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 1000 });
await page.goto(url, { waitUntil: 'networkidle0', timeout: 60000 });
await page.screenshot({ path: 'page-full.png', fullPage: true });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
await browser.close();
}
Run it with node archive.mjs https://example.com. Puppeteer’s screenshot API documents screenshot options. Its PDF API documents that PDF generation uses print CSS by default. To request screen media before creating a PDF, add this before page.pdf():
await page.emulateMediaType('screen');
Print and screen styles can differ: page breaks, hidden elements, colors, and layout may change. Inspect representative PDFs from the pages you archive.
5. Make the archive repeatable
- Decide what the record should include: visible viewport, entire scrollable page, or a specific element.
- Choose a readiness condition based on the content you need. Prefer a page-specific selector when a key element must appear; a timeout alone does not confirm completeness.
- Save the screenshot and PDF together. Use a sanitized title and date, for example
2026-10-04-example-page.pngand2026-10-04-example-page.pdf. - Record the source URL and capture time in a small manifest or adjacent text file. This is a practical provenance convention, not an archival standard.
- Open the outputs and check image extent, PDF pagination, colors, and missing content. Back up archived files if you need to retain them.
For multiple URLs, put the URLs in a list and loop through them with Playwright or Puppeteer. Use a stable filename derived from the URL or page title, handle each page’s navigation failure separately, and close the browser in a finally block so an error does not leave it running.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Make one GET request to capture a URL as an image or PDF; see the ScreenshotNeo API documentation for request options and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. ScreenshotNeo supports PDF output as well as PNG, JPEG, and WebP. See ScreenshotNeo for product details and the docs for PDF request parameters.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
| The screenshot is blank or incomplete. | The page had not rendered the needed content before capture, or content loads only after scrolling. | Wait for a meaningful selector or page-specific state. For lazy-loaded content, scroll through the page before a full-page capture. |
| The command finishes but the page was still loading. | Chrome’s timeout is a maximum wait; capture can happen before loading finishes. | Use Playwright or Puppeteer with an explicit readiness condition, or adjust the CLI timeout while checking the resulting output. |
| The PDF looks different from the browser. | PDF generation follows print styles by default. | Inspect print CSS, or emulate screen media before generating the PDF when screen styling is required. |
| The full page is missing below the fold. | A viewport screenshot was captured. | Set fullPage: true in the Playwright or Puppeteer screenshot call. |
| Navigation times out on a dynamic site. | The page may keep requests active, or the chosen wait condition may not match its behavior. | Use a more suitable navigation wait and then wait for the specific content needed. Do not treat a longer timeout as evidence of completeness. |
| Output files are overwritten. | Every run uses the same path. | Include a date and sanitized title or another stable unique identifier in each filename. |
Performance, reliability, and cost
The reviewed documentation provides no comparative capture benchmarks, so there is no supported speed ranking between these methods. Capture time depends on navigation, page behavior, readiness conditions, and the work your script performs. Keep each page’s timeout bounded, avoid launching unnecessary browser instances in a batch job, and record failures so one bad URL does not hide successful captures.
Browser automation runs on infrastructure you operate, so account for browser installation, runtime, and storage for the resulting files. The documented CLI and libraries do not provide an archiving guarantee: preserve both outputs and provenance metadata in the storage system your workflow already uses. With ScreenshotNeo, only clean shots are billed, and the response includes verdict and billing headers; see its current plan details at screenshotneo.com.
FAQ
Is a screenshot a complete copy of the webpage?
No. It records a visual rendering at capture time. It does not preserve all interactive behavior or guarantee access to the underlying page resources later.
Should I keep both the screenshot and PDF?
Keep both when you need a fixed visual image and a document that is convenient to reopen or share. Check each output because their layouts and content can differ.
Can I archive pages on a schedule?
Yes. Run a script on your scheduler, feed it a URL list, and use unique filenames plus a manifest with the URL and capture time. The browser libraries provide capture operations; scheduling and retention are part of your environment.


