ScreenshotNeo

BlogHTML to image & PDF

How to Save a Webpage as a PDF with Puppeteer

Use Puppeteer’s page.pdf() to turn a webpage into a PDF. Learn setup, print styling, page sizing, output options, troubleshooting, and a one-call alternative.

By the ScreenshotNeo team4 October 202611 min read

To save a webpage as a PDF with Puppeteer, navigate to it in Chromium and call page.pdf(). Set path to write a file, and use printBackground: true if the document needs background colors or images. Puppeteer renders PDFs with print CSS by default. See the official PDF generation guide and the PDFOptions reference.

Install Puppeteer

Start with a Node.js project and install Puppeteer. The package normally downloads a compatible browser during installation; if your environment manages Chrome separately, follow Puppeteer’s current installation and configuration guidance.

mkdir webpage-to-pdf
cd webpage-to-pdf
npm init -y
npm install puppeteer

Save the following as save-page.mjs. The .mjs extension enables ES module imports without changing the project configuration.

Save a webpage to a PDF

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com';
const outputPath = process.argv[3] ?? 'page.pdf';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle2' });
  await page.pdf({ path: outputPath });
  console.log(`Saved ${url} to ${outputPath}`);
} finally {
  await browser.close();
}

Run it with:

node save-page.mjs https://example.com example.pdf

The navigation wait shown here follows Puppeteer’s guide example. It is not a universal readiness rule: pages that keep network connections open or render content after navigation may need a different wait condition or an explicit page-specific readiness check. Close the browser in a finally block so it is released when navigation or PDF generation throws.

Choose the right readiness condition

The PDF can only contain what the page has rendered by the time PDF generation starts. Choose a wait strategy based on how the target site loads:

Approach Use it when Trade-off
page.goto(url) The page is mostly server-rendered and the initial load event is sufficient. Client-side content may still be loading.
waitUntil: 'networkidle2' The page settles after a short period with few active network connections. Analytics, polling, and long-lived requests can delay or prevent network quiet.
page.waitForSelector() A known element indicates that the content you need has appeared. The element may exist before its data or images are complete.
page.waitForFunction() Your app exposes a state that signals the report or view is ready. Requires a meaningful readiness condition from the page.

For a page that renders a report after an API request, wait for a selector or application state rather than assuming that network quiet means the report is complete. For example:

await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 15000 });
await page.pdf({ path: 'report.pdf', printBackground: true });

Replace the example selector with one that actually represents readiness on the target page. A timeout helps the job fail visibly instead of silently producing an incomplete document.

Configure paper, margins, media, and print styling

page.pdf() uses the print CSS media type. If the PDF should resemble the on-screen design, call page.emulateMediaType('screen') first. Print styles can intentionally hide navigation, change layout, and adjust colors, so inspect the output when the result differs from the browser view. [Puppeteer Page API]

Common PDF options

Option What it controls Documented default or behavior
path Output filename. Omitted means no disk write; relative paths resolve from the process working directory.
format Named paper format, such as A4 or Letter. Defaults to Letter; takes priority over width and height.
width, height Paper dimensions as numbers or strings with units. Use these for a custom paper size; format takes priority if also supplied.
landscape Landscape orientation. false.
margin Top, right, bottom, and left page margins. No margins are set unless specified.
printBackground Include CSS background graphics. false.
preferCSSPageSize Prioritize page dimensions from CSS @page. false; content is scaled to fit the requested paper size.
pageRanges Select pages such as 1-3, 7. Empty string prints all pages.
scale Scale page rendering. 1; accepted range is 0.1–2.
displayHeaderFooter Show print header and footer templates. false.
waitForFonts Wait for document.fonts.ready. true; a background page may need page.bringToFront().
timeout PDF generation timeout in milliseconds. 30,000 ms; 0 disables this timeout.
omitBackground Hide the default white page background, allowing transparency. false.
outline, tagged Generate a document outline or tagged PDF. Both are experimental options in the API reference; verify support and output in your installed version.

Paper size and margins

Use a named format for standard paper, or set a custom width and height. When a page defines its own paper dimensions in CSS, preferCSSPageSize: true makes those dimensions take priority. The official PaperFormat reference lists supported paper names.

await page.pdf({
  path: 'invoice.pdf',
  format: 'A4',
  landscape: false,
  margin: {
    top: '12mm',
    right: '14mm',
    bottom: '16mm',
    left: '14mm',
  },
});

For a page controlled by CSS @page rules:

await page.pdf({
  path: 'css-sized.pdf',
  preferCSSPageSize: true,
  printBackground: true,
});

When format and dimensions compete, format takes priority over width and height. When preferCSSPageSize is true, CSS page size takes priority over those options. Avoid setting conflicting sizing rules unless you have a specific reason.

Background graphics are off by default. Enable them when important design elements use background fills or images:

await page.pdf({ path: 'designed-report.pdf', printBackground: true });

Browsers may adjust colors for print output. To request more faithful CSS colors, add this to the page’s print stylesheet:

@media print {
  html {
    -webkit-print-color-adjust: exact;
  }
}

Screen media instead of print media

If the website has no useful print layout and you want its screen styles, emulate screen media before calling page.pdf():

await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-layout.pdf', printBackground: true });

Screen media does not itself guarantee that the layout fits paper well. Wide content may be scaled or split across pages, so choose page dimensions and inspect the PDF.

Headers, footers, and page ranges

Enable displayHeaderFooter to print header and footer templates. Puppeteer recognizes template classes for date, title, URL, page number, and total pages:

await page.pdf({
  path: 'selected-pages.pdf',
  displayHeaderFooter: true,
  headerTemplate: '<span class="title"></span>',
  footerTemplate: '<span class="pageNumber"></span> / <span class="totalPages"></span>',
  pageRanges: '1-5, 8',
});

Templates are HTML, but their rendering constraints are narrower than a full webpage. Keep them simple, and allow enough margin so the header and footer do not overlap body content. pageRanges uses a string such as 1-5, 8, 11-13; an empty string means all pages.

Return PDF bytes instead of writing a file

When a server needs to send the PDF in an HTTP response or pass it to another function, omit path and use the returned bytes. The Page API documents page.pdf() as returning a Uint8Array.

const pdfBytes = await page.pdf({
  format: 'A4',
  printBackground: true,
});

// In a Node.js HTTP handler, for example:
response.setHeader('Content-Type', 'application/pdf');
response.setHeader('Content-Disposition', 'attachment; filename="page.pdf"');
response.end(Buffer.from(pdfBytes));

If you need a filesystem path, supply path; if you need in-memory bytes, omit it. Be mindful that holding large PDFs in memory increases the memory used by your process.

Complete example with practical settings

This version waits for a known content marker, emits a styled A4 PDF, and guarantees browser cleanup. Adapt the marker, URL, and print styles to the site you control.

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch();

try {
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(30000);

  await page.goto(url, { waitUntil: 'domcontentloaded' });
  // Replace this selector with a marker that indicates the content is ready.
  await page.waitForSelector('main', { timeout: 15000 });

  await page.pdf({
    path: 'page.pdf',
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
    margin: { top: '12mm', right: '12mm', bottom: '14mm', left: '12mm' },
    waitForFonts: true,
    timeout: 30000,
  });
} finally {
  await browser.close();
}

The main selector is only an example; many pages have one before the useful content finishes loading. For dynamic pages, choose an application-specific ready state or wait for a response that provides the data you need. This example documents an approach, not a guarantee for every site or deployment environment.

cURL, Python, and Node.js alternatives

Puppeteer is a Node.js library, so its direct workflow is JavaScript. These snippets use ScreenshotNeo’s hosted screenshot API for developers who want the rendered page returned by one HTTP request instead of managing a local browser. The API can return an image or PDF; for a PDF response, request PDF output using the documented API option names in the ScreenshotNeo API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o page.pdf

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://example.com",
        "format": "pdf",
    },
    timeout=90,
)
r.raise_for_status()
with open("page.pdf", "wb") as output:
    output.write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
  format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot API returned ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
  writeFile('page.pdf', Buffer.from(await res.arrayBuffer()))
);

These API examples are separate from Puppeteer’s browser workflow. Check the docs for the current PDF parameter and response behavior before using them; do not expose an API key in client-side code.

Or skip the browser setup

ScreenshotNeo returns a screenshot or PDF from one API request, so you do not have to install and operate Puppeteer and Chromium for this capture. Cookie banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients the take_screenshot, get_page_info, and capture_pdf tools.

One-call PDF example (see the ScreenshotNeo docs for the PDF option name and all parameters):

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -d format=pdf \
  -o page.pdf

ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Create a free account and get 1,000 screenshots a month with no card.

Troubleshooting Puppeteer PDF output

Symptom Likely cause Fix
PDF is blank or missing page content PDF generation began before client-rendered content appeared, or the site returned an interstitial. Wait for a content-specific selector or application ready state; inspect the final page URL and page content before printing.
Navigation hangs or times out The page keeps network requests active, or the selected navigation condition is too strict for that site. Use a less restrictive navigation condition such as domcontentloaded, then wait for the actual content marker. Set an intentional navigation timeout.
Background colors or images are absent printBackground defaults to false. Set printBackground: true.
PDF layout differs from the browser page.pdf() uses print CSS by default. Review the site’s @media print styles, or call page.emulateMediaType('screen') before generating.
Page size is wrong A configured format, dimensions, or CSS @page rule is taking priority. Choose one sizing strategy: named format, explicit width and height, or CSS sizing with preferCSSPageSize: true.
Text or images are clipped Content exceeds the paper width, or margins and scale do not suit the page. Try landscape, a suitable paper size, smaller margins, or a scale between 0.1 and 2; check wide tables and fixed-width elements.
Custom fonts are missing Fonts have not loaded, or the page is in a background tab while Puppeteer waits for them. Keep the default waitForFonts: true; if needed, call page.bringToFront() and verify the font resource loads.
Colors look washed out Print rendering adjusts colors by default. Use -webkit-print-color-adjust: exact in the page’s print CSS and include backgrounds if required.
page.pdf() does not create a file The path option was omitted, or a relative path points somewhere unexpected. Set path explicitly; resolve relative paths from the process working directory or use an absolute path.
PDF generation times out Fonts or rendering have not completed within the configured timeout. Check whether the page is still progressing, set a suitable timeout, and fix stalled resources. The documented default is 30 seconds; 0 disables the PDF timeout.

Performance, reliability, and cost considerations

  • Wait only for what matters. A full network-idle wait can be slow on pages with ongoing requests. A specific content selector often gives a clearer readiness signal, when the page exposes one.
  • Reuse browser processes thoughtfully. For a one-off script, launch and close one browser as shown. For a service handling many jobs, browser startup and memory use become operational concerns; bound concurrency and close pages and browsers on error. The supplied Puppeteer references do not prescribe a production deployment architecture.
  • Keep PDF size and memory in view. Printing backgrounds, large images, and long pages can increase render time and output size. Returning bytes instead of writing a file also keeps those bytes in application memory until they are sent or released.
  • Set bounded timeouts. Navigation and PDF generation are separate operations with separate failure points. Log which stage failed and the target URL so retries can be limited to transient failures.
  • Costs depend on where Chromium runs. Puppeteer itself is the browser automation library; infrastructure, compute, storage, and maintenance costs depend on your environment. No deployment or cost benchmark is established by the cited Puppeteer documentation.
  • Hosted alternative pricing. ScreenshotNeo offers 1,000 shots each month free with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Confirm current details on its site.

Frequently asked questions

Does Puppeteer create a PDF from the whole page?

page.pdf() prints the page according to its print layout and paper configuration. Use print CSS, paper sizing, and page ranges to shape the result; inspect long or responsive pages for overflow.

Can I save only selected pages?

Yes. Set pageRanges, for example '1-3, 7'. Leave it empty to print all pages.

Can I use CSS @page rules?

Yes. Set preferCSSPageSize: true to prioritize CSS page dimensions over the PDF width, height, or format settings.

Can the PDF have a transparent background?

The PDFOptions reference provides omitBackground, which hides the default white background and allows PDF transparency. Check how the target page’s own background styles affect the result.

Which Puppeteer version should I use?

Use the version installed in your project and consult its matching API reference. The dossier’s current-reference snapshot identified version 25.12.0; Puppeteer’s APIs can change over time.

References