ScreenshotNeo

BlogHTML to image & PDF

Convert a URL to PDF in Node.js Using Puppeteer

Render a web page to PDF with Puppeteer in Node.js. Learn setup, print and screen layouts, page options, readiness checks, troubleshooting, and a no-browser alternative.

By the ScreenshotNeo team4 October 20269 min read

Use Puppeteer to launch a browser, navigate to a fully qualified URL, and call page.pdf(). Puppeteer prints with the page’s print CSS by default; set options such as format and printBackground to control the output. Close the browser in a finally block so it is also closed when navigation or PDF generation fails.

This guide uses the Puppeteer API documented in version 25.12.0. Puppeteer is guaranteed to work with its bundled browser; using another browser is at your own risk. See the official Getting started guide, Page API, and PDF options reference.

1. Install Puppeteer

In a new Node.js project, install Puppeteer:

npm init -y
npm install puppeteer

Puppeteer downloads a compatible browser as part of its installation. If your project uses ES module imports, set "type": "module" in package.json, or save the script with an .mjs extension.

2. Convert a URL to PDF

Save this as url-to-pdf.mjs. Pass the URL as the first command-line argument and optionally an output path as the second:

import puppeteer from 'puppeteer';

const url = process.argv[2];
const outputPath = process.argv[3] ?? 'page.pdf';

if (!url) {
  console.error('Usage: node url-to-pdf.mjs <https://example.com> [output.pdf]');
  process.exit(1);
}

let parsedUrl;
try {
  parsedUrl = new URL(url);
} catch {
  console.error('Provide a valid, fully qualified URL, such as https://example.com');
  process.exit(1);
}

if (!['http:', 'https:'].includes(parsedUrl.protocol)) {
  console.error('Only HTTP and HTTPS URLs are supported by this script.');
  process.exit(1);
}

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(45_000);

  const response = await page.goto(parsedUrl.href, {
    waitUntil: 'load',
    timeout: 45_000,
  });

  if (!response) {
    throw new Error('Navigation did not return a main-resource response.');
  }
  if (!response.ok()) {
    throw new Error(`Page returned HTTP ${response.status()} ${response.statusText()}`);
  }

  await page.pdf({
    path: outputPath,
    format: 'A4',
    printBackground: true,
    waitForFonts: true,
  });

  console.log(`Saved ${outputPath}`);
} finally {
  await browser.close();
}

Run it with:

node url-to-pdf.mjs https://example.com example.pdf

The URL must include its scheme, such as https://. page.goto() resolves with the main resource response, including the response after redirects; a resolved navigation does not by itself mean the HTTP status is successful. The script checks the response status before writing the PDF. The output path is relative to the current working directory unless you provide an absolute path.

3. Choose when the page is ready

The waitUntil option controls which navigation lifecycle event Puppeteer waits for. Its default is load. You can use one condition or an array, in which case every listed condition must fire. No single readiness condition suits every site.

Condition Use when Trade-off
load The page’s load event is a reasonable signal that its main document and dependent resources have finished loading. Some pages render important content later in application code.
networkidle0 or networkidle2 You want to wait for network activity to become idle, and the site’s requests eventually settle. Analytics, polling, streaming, or other ongoing requests can make network-idle waiting unsuitable.
Page-specific signal The site renders content after navigation and exposes a known selector or application-ready condition. You need to identify a signal that actually means the content you want is ready.

For a page-specific selector, navigate first and then wait for the content needed in the PDF:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 15_000 });
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

Replace the selector with one that exists on your target page. If your app exposes a different reliable ready signal, use that instead. Navigation timeouts and selector timeouts are separate waits and should be chosen for the site’s expected behavior.

See Puppeteer’s navigation reference and wait options.

4. Set the PDF layout

page.pdf() uses print CSS media by default. This often produces a document-oriented layout, which may differ from what a browser displays on screen. To use screen media instead, call page.emulateMediaType('screen') before generating the PDF.

// Print layout (the default):
await page.pdf({ path: 'print-layout.pdf', format: 'A4', printBackground: true });

// Screen CSS layout:
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-layout.pdf', format: 'A4', printBackground: true });

When the site’s colors or backgrounds matter, set printBackground: true. Print rendering may otherwise alter colors, and background graphics are not enabled by default.

Paper size, margins, and page breaks

The documented PDF options include:

Option What it controls
format Named paper size. The documented default is letter; common values include A4 and Letter.
landscape Whether the paper is landscape rather than portrait.
margin Page margins.
scale Scale applied to the rendered page.
pageRanges Which pages to include in the PDF.
preferCSSPageSize Whether CSS page size takes priority over PDF paper dimensions. The documented default is false.
path Where to write the PDF. Relative paths resolve from the current working directory.
waitForFonts Whether to wait for fonts before printing. The documented default is true.

For a layout with CSS-defined paper dimensions, add an @page rule and set preferCSSPageSize: true:

await page.pdf({
  path: 'custom-page-size.pdf',
  preferCSSPageSize: true,
  printBackground: true,
  margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
});
/* Include this in the page's print stylesheet. */
@page {
  size: A4 landscape;
  margin: 12mm;
}

If you prefer to specify the paper in the PDF call, use format or the documented width and height options, and leave CSS page sizing at its default. The full option definitions are in the PDFOptions reference.

5. Handle common input and rendering cases

Authenticated pages

If the target page requires a login, a fresh Puppeteer page will not automatically have that site’s session. Establish the required session in the browser context before navigation, using an authentication flow appropriate to your application. Do not put credentials into a URL or commit secrets into the script.

HTML already in your script

For HTML you already have, page.setContent(html) sets the page content. This is useful when rendering supplied markup; it is not a replacement for URL navigation when the remote page and its resources need to load.

const page = await browser.newPage();
await page.setContent('<main><h1>Report</h1><p>Generated content</p></main>');
await page.pdf({ path: 'generated.pdf', format: 'A4', printBackground: true });

Relative images, stylesheets, and fonts in supplied HTML need valid resource URLs or another resource-loading approach. The setContent reference documents the method and its optional wait parameters.

URLs that return a PDF

This workflow is for rendering a web page to a PDF. In headless shell mode, Puppeteer does not support navigating to a PDF document with page.goto(). If the input URL itself serves a PDF, handle that document as a PDF download or use a suitable PDF-processing workflow instead of treating it as an HTML page to print.

Page content and page breaks

For long pages, check the rendered PDF for clipped content, awkward breaks, and missing backgrounds. Page-break behavior is affected by the page’s print CSS. If you control the site, use print-specific styles and CSS page rules to make the intended document layout explicit.

6. Troubleshoot conversion failures

Symptom Likely cause Fix
Invalid URL error The URL is malformed or lacks a scheme. Use a fully qualified URL such as https://example.com.
Navigation timeout The page did not reach the selected lifecycle condition before the timeout, or it keeps making requests. Choose a lifecycle condition that fits the page, set a suitable timeout, or wait for a page-specific ready selector after navigation.
Navigation fails with a certificate or network error The server is unreachable, the SSL connection fails, or the main resource cannot load. Check that the URL is reachable from the machine running Node.js and that its certificate and network path are valid.
The script reports an HTTP error status The server returned an error response even though navigation produced a response. Check the status and target URL. Handle redirects, access restrictions, and server errors according to the site’s requirements.
PDF is missing late-loaded content The navigation event fired before the application rendered that content. Wait for a known page-specific selector or readiness signal before calling page.pdf().
PDF colors or background blocks are missing Print rendering changes colors, or print backgrounds are disabled. Set printBackground: true; use screen media only when that layout is what you want.
PDF layout differs from the browser page.pdf() uses print media by default. Use print CSS for a document layout, or call page.emulateMediaType('screen') before printing to use screen CSS.
Wrong paper dimensions The PDF paper option and the page’s CSS page size do not agree. Choose a paper format explicitly, or define @page dimensions and set preferCSSPageSize: true.
Font is absent or substituted The font was not available to the page when it was printed. Ensure the page can load the required font and keep waitForFonts: true, the documented default.
Browser process remains open after an error The script does not close the browser on every code path. Put browser.close() in a finally block.
Cannot navigate to a PDF URL The target is a PDF document and the run uses headless shell mode. Use a PDF download or PDF-processing path for PDF input rather than the web-page printing flow.

7. Performance, reliability, and cost

Each conversion needs a browser page to load the target and render it. The source documentation does not provide a universal conversion-time or memory benchmark, so measure representative pages in the environment where your script will run. Large pages, slow external assets, and readiness conditions that wait on ongoing network activity can extend the job.

  • Set navigation and readiness timeouts so a stalled page does not wait indefinitely.
  • Use a page-specific signal when a site’s meaningful content appears after the ordinary load event.
  • Close the browser in finally to release browser resources on both success and failure.
  • Use Puppeteer’s bundled browser for the documented compatibility guarantee; another browser is at your own risk.
  • Account for the time and compute used to run Node.js and the browser in your own environment. Puppeteer’s cited API references do not specify a conversion price.

Or skip the browser setup

If you only need to capture a page as a PDF, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call endpoint can return a PDF; see the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses include X-Page-Verdict and X-Billed headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Does Puppeteer save a PDF from a URL directly?

Not in one call: navigate to the page with page.goto(), then generate the PDF with page.pdf().

Why does the result use a different layout than the live page?

PDF generation uses print CSS by default. Emulate screen media before printing if the screen stylesheet is the desired layout.

Can I choose which pages appear in the output?

Yes. The PDF options include pageRanges; consult the current PDFOptions reference for its accepted syntax.

Which Puppeteer browser should I use?

Puppeteer guarantees compatibility with its bundled browser. Using another browser is at your own risk.

References