How to Scrape Web Pages with Puppeteer and Save Them as PDFs
Use Puppeteer to load a web page, wait for the content you need, and save it as a PDF. Configure print styles, paper size, margins, and page ranges.
To save a web page as a PDF with Puppeteer, launch a browser, navigate to the page, wait for the content you need, and call page.pdf(). Puppeteer prints using print CSS by default. The example below waits for a page-specific element, checks the main HTTP response, enables background graphics, and closes the browser even if a step fails.
1. Install Puppeteer
Use Node.js with an ES module project. Puppeteer downloads a compatible browser during installation by default.
mkdir puppeteer-pdf
cd puppeteer-pdf
npm init -y
npm pkg set type=module
npm install puppeteer
The code below uses the Puppeteer 25.12.0 API documented in the [official API reference](https://pptr.dev/api). If you need repeatable deployments, pin the version in your lockfile and install dependencies with your package manager’s frozen-lockfile option.
2. Navigate, wait for content, and create the PDF
Save this as save-page.mjs. Pass the target URL as the first command-line argument. Replace main with a selector or condition that reflects when the content you intend to print is ready.
import puppeteer from 'puppeteer';
const url = process.argv[2];
if (!url) {
console.error('Usage: node save-page.mjs https://example.com');
process.exit(1);
}
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
page.setDefaultNavigationTimeout(45_000);
page.setDefaultTimeout(15_000);
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (response && !response.ok()) {
throw new Error(`Main document returned HTTP ${response.status()}`);
}
// Wait for the page-specific content that must appear in the PDF.
await page.locator('main').wait();
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
waitForFonts: true,
timeout: 30_000,
});
console.log('Saved page.pdf');
} finally {
await browser.close();
}
Run it with:
node save-page.mjs https://example.com
This is a runnable starting point, but the readiness selector must match the target page. Puppeteer’s navigation API returns the main resource response after redirects; it can return null for cases such as about:blank or same-page hash navigation. A completed navigation alone does not prove a successful HTTP status or that a client-rendered page has finished filling in its content.
3. Choose a reliable readiness condition
Prefer a signal tied to the content being printed over an arbitrary sleep. Puppeteer locators wait for an element and its relevant state. For a custom application condition, use page.waitForFunction().
Wait for a selector
await page.locator('[data-report-ready="true"]').wait();
Use a stable selector that appears only when the report or article content is available. A generic container such as main may exist before its data has loaded.
Wait for an application condition
await page.waitForFunction(() => {
const report = document.querySelector('#report');
return report && report.textContent.trim().length > 0;
}, { timeout: 15_000 });
The function runs in the page context and is polled until it returns a truthy value or times out. Pass only conditions that can be evaluated in the browser page.
When navigation wait conditions help
page.goto() accepts wait conditions such as domcontentloaded, load, and networkidle0 or networkidle2. A stricter network-idle condition can be unsuitable for pages that keep analytics, polling, or streaming requests open. Even a successful load event may happen before an application renders the content you need. Use navigation waits to establish a baseline, then wait for the page-specific condition.
4. Configure the PDF
Puppeteer’s PDFOptions reference lists the output and rendering controls. The options below are the ones developers most often need to make explicit.
| Option | What it controls | When to use it |
|---|---|---|
path |
Output filename. Relative paths resolve from the process working directory. | Set it to write a file; omit it if you want the PDF bytes returned instead. |
format |
Paper preset such as A4 or Letter. |
Use a named paper size for predictable document output. |
width, height |
Explicit paper dimensions. | Use dimensions when a standard preset is not appropriate. Follow Puppeteer’s accepted units. |
landscape |
Orientation; defaults to portrait. | Set true for wide tables or charts. |
margin |
Page margins. | Specify top, right, bottom, and left values when the default layout clips or crowds content. |
printBackground |
Whether background graphics are printed; defaults to false. |
Set true when colored blocks or background images matter. |
scale |
Scales the page rendering. | Adjust only when content needs to fit; excessive scaling can make text hard to read. |
pageRanges |
Which pages to include. | Use ranges such as 1-3 when only part of a long document is needed. |
preferCSSPageSize |
Whether CSS @page dimensions take priority over PDF dimensions; defaults to false. |
Set true when the site defines its intended paper size in CSS. |
waitForFonts |
Waits for fonts before printing; documented default is true. |
Keep enabled when font metrics affect line breaks and pagination. |
timeout |
Maximum time for PDF generation. | Set a bound appropriate to the document and execution environment. |
For example, use US Letter, landscape, explicit margins, and a page range like this:
await page.pdf({
path: 'report.pdf',
format: 'Letter',
landscape: true,
margin: { top: '12mm', right: '10mm', bottom: '12mm', left: '10mm' },
printBackground: true,
pageRanges: '1-5',
timeout: 30_000,
});
Choose either the named format or explicit dimensions to express the paper size. When the site has deliberate @page rules, preferCSSPageSize: true lets those rules take priority.
5. Control print styling and pagination
page.pdf() uses print media, so the page’s print stylesheet may hide navigation, change colors, or alter layout. This is usually desirable for documents. To render the screen stylesheet instead, set screen media before creating the PDF:
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-layout.pdf', format: 'A4', printBackground: true });
Use the default print media when you want the site’s print design. Use screen media only when the screen layout is the intended output. Review both the first page and page breaks when the site contains long tables, fixed headers, or large images. Background colors and images are omitted unless printBackground is enabled.
6. Add authentication or interact before printing
For a page you are authorized to access, set cookies or headers before navigation, or perform the required interaction before generating the PDF. Keep secrets out of source control and logs.
await page.setExtraHTTPHeaders({ Authorization: `Bearer ${process.env.PAGE_TOKEN}` });
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.locator('#report').wait();
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });
For cookie-based sessions, use page.setCookie(...) before navigating. If printing requires clicking a control such as “Show details,” use a locator action and then wait for the resulting content. Do not attempt to bypass access controls. Robots rules are crawler guidance and do not grant access authorization; check applicable site terms and access restrictions separately, as explained in IETF RFC 9309.
7. Capture many URLs without leaking browser resources
For a small batch, reuse one browser and create a fresh page for each URL. Close each page in a finally block. A single long-lived page can carry cookies, local storage, or application state from one capture into the next.
import puppeteer from 'puppeteer';
const urls = ['https://example.com', 'https://example.org'];
const browser = await puppeteer.launch();
try {
for (const [index, url] of urls.entries()) {
const page = await browser.newPage();
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
if (response && !response.ok()) {
throw new Error(`${url}: HTTP ${response.status()}`);
}
await page.locator('main').wait();
await page.pdf({ path: `page-${index + 1}.pdf`, format: 'A4', printBackground: true });
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
For production batches, limit concurrency to what the machine can support, retry only transient failures with a cap, and record the URL and failure stage for each item. Avoid unlimited parallel page creation: every open page consumes browser and system resources. Do not retry a deterministic failure such as a missing selector indefinitely.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its API can return screenshots or PDFs; see the API documentation for PDF options and parameters. For a one-call image example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
These examples save an image. Use the documented PDF parameters when you need PDF output. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF is blank or missing page content | The app had not rendered the content when printing began, or the selector was too broad. | Wait for a page-specific selector or a waitForFunction condition tied to the required data. |
page.goto() returns a 404 or 500 response without throwing |
Navigation completion and HTTP success are separate checks. | Inspect response.status() or response.ok() and handle the status explicitly. |
| Navigation timeout | The chosen lifecycle event did not occur in time, or the site keeps network requests open. | Choose an appropriate waitUntil condition, set a deliberate navigation timeout, and wait separately for the content you need. |
| Locator timeout | The selector is wrong, content is absent, or the page has not reached the expected state. | Confirm the selector in the page DOM, check authorization and app state, and set a suitable locator timeout. |
| Colors or background images are missing | printBackground is false by default. |
Set printBackground: true. |
| Layout differs from the browser screenshot | PDF output uses print media by default. | Keep print CSS if intended, or call page.emulateMediaType('screen') before printing. |
| Content is clipped or unexpectedly scaled | Paper dimensions, CSS @page, margins, or scale conflict. |
Choose a paper size, inspect the site’s print CSS, adjust margins, and set preferCSSPageSize intentionally. |
| Fonts or line wrapping differ | Fonts may not have loaded or may not be available to the browser. | Keep waitForFonts: true, verify font requests and availability, then review pagination. |
| Browser process remains after an error | Cleanup did not run after a thrown error. | Put browser.close() in a finally block. |
| Navigation response is null | Some navigations, including about:blank and same-document hash changes, have no main resource response. |
Handle the null case instead of dereferencing it. |
10. Performance, reliability, and cost
- Wait narrowly. A selector or application condition can avoid waiting for unrelated background requests. Fixed sleeps add delay and still do not guarantee readiness.
- Reuse the browser process for batches. Launching once and closing each page reduces repeated browser startup work while keeping individual page state isolated.
- Bound time and concurrency. Set navigation, selector, and PDF timeouts. Limit simultaneous pages to avoid exhausting memory or CPU; tune based on the workload and available machine rather than assuming a universal rate.
- Use stable inputs. Dynamic ads, personalization, changing page data, and remote fonts can change pagination. For repeatability, control authentication, viewport where relevant, page state, and capture timing.
- Budget for browser operations. Puppeteer itself is an open-source automation library, but running a browser consumes compute and memory. Infrastructure, storage, retries, and operational maintenance determine the cost; there is no universal per-PDF price or runtime established here.
- Keep failures observable. Log navigation status, URL, wait condition, and whether PDF generation failed. Retry transient network failures selectively, and avoid retrying invalid URLs or selectors as if they were transient.
11. FAQ
Can Puppeteer scrape the page text as well as create a PDF?
Yes. Puppeteer can inspect the DOM and extract text before or alongside PDF generation. Keep extraction separate from PDF layout when the two tasks have different readiness requirements.
Does page.pdf() print exactly what I see on screen?
Not by default. It renders print media. Use screen media emulation if the screen stylesheet is the required output.
Can I save the PDF to a buffer instead of a file?
Yes. Omit path from the PDF options and use the returned PDF data in your application; see the Page API.
Does robots.txt give permission to scrape a page?
No. RFC 9309 says robots rules are not access authorization. Check site terms and other access restrictions separately.


