ScreenshotNeo

BlogHTML to image & PDF

How to Generate a PDF from a URL

Save a webpage as a PDF manually, from the command line with Chrome Headless, or in JavaScript with Puppeteer. Control print layout, timing, and output.

By the ScreenshotNeo team29 September 202610 min read

How to Generate a PDF from a URL

To generate a PDF from a URL, use your browser’s print-to-PDF feature for a one-off save, Chrome Headless for a command-line job, or Puppeteer for repeatable JavaScript automation. For example, Chrome can save a URL as output.pdf with chrome --headless --print-to-pdf https://example.com/. Puppeteer gives you more control over when the page is ready and how the PDF is laid out.

The result depends on the page’s print styles, scripts, fonts, authentication, and resources. A successful navigation does not guarantee that every dynamic element has appeared or that the page will paginate as you expect. This guide covers the three approaches, runnable examples, layout controls, common failures, and ways to make automated output more reliable.

1. Choose the right way to save the page

Approach Best for What you control
Browser print dialog A single manual save Print preview and the browser’s available print settings
Chrome Headless A simple command-line conversion PDF output, generated headers and footers, and timing flags
Puppeteer Recurring jobs or an application workflow Navigation waits, JavaScript execution, print media, paper, margins, and page ranges
ScreenshotNeo A hosted API request that returns a PDF API capture options, without setting up a local browser

For a single page, open it in a browser and use its print-to-PDF flow. The exact menu names and steps vary by browser, operating system, and version; check the print preview before saving. This is convenient, but it is not a repeatable automation interface.

For a scriptable process, Chrome Headless is the shortest path. Choose Puppeteer when you need code to set readiness conditions or PDF layout options. If you need a hosted conversion endpoint, see the ScreenshotNeo option below.

2. Generate a PDF with Chrome Headless

Chrome’s headless command line can print a URL to a PDF in the current working directory. Install Chrome or Chromium and make sure the executable is available as chrome in your shell; otherwise use the full executable path for your platform.

chrome --headless --print-to-pdf https://example.com/

The output is named output.pdf by default. Run the command from a directory where you can write that file. To avoid the generated print header and footer, which include the date and time, URL, and page number, use --no-pdf-header-footer:

chrome --headless --no-pdf-header-footer --print-to-pdf https://example.com/

Some older Chrome versions used --print-to-pdf-no-header for this behavior. If Chrome rejects the newer flag, check the command-line reference for the version you have installed and try the legacy flag where applicable.

For pages that need more time before printing, Chrome documents --timeout as a maximum wait before capture, even if the page is still loading. --virtual-time-budget fast-forwards timer-driven page code. These options help with some delayed content, but neither guarantees that every asynchronous request, image, or application state is complete.

chrome --headless --timeout=10000 --print-to-pdf https://example.com/

Use the timeout as a limit, not as proof of readiness. If a page displays its main content only after an API call or a user action, inspect the page behavior and choose a workflow that can wait for the relevant condition. Chrome’s documented options are listed in the Chrome Headless command-line reference.

3. Generate a PDF with Puppeteer

Puppeteer launches a browser controlled from JavaScript. Install it in a Node.js project, save this as url-to-pdf.js, and run it with Node. The example navigates to the target, waits for a network-idle condition, writes a PDF, and closes the browser even if generation fails.

The browser renders a URL using page and print styles, then writes the resulting pages into a PDF.
The browser renders a URL using page and print styles, then writes the resulting pages into a PDF.
const puppeteer = require('puppeteer');

async function main() {
  const url = process.argv[2] || 'https://example.com/';
  const browser = await puppeteer.launch();

  try {
    const page = await browser.newPage();
    await page.goto(url, {
      waitUntil: 'networkidle2',
      timeout: 60000,
    });
    await page.pdf({ path: 'output.pdf' });
  } finally {
    await browser.close();
  }
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

Install the dependency in the project with npm install puppeteer, then run node url-to-pdf.js https://example.com/. Puppeteer’s official guide documents page.pdf() as the PDF generation method and shows the launch, navigation, and save pattern. See the Puppeteer PDF generation guide.

Wait for the condition the page actually needs

waitUntil: 'networkidle2' is one navigation condition, not a universal signal that a page is visually complete. Sites with polling, analytics, streaming requests, lazy-loaded images, or delayed UI may behave differently. Conversely, waiting for all network activity to stop can take too long on pages that keep requests open.

When you know a specific element signals that the page is ready, wait for it explicitly after navigation:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
await page.waitForSelector('main article', { timeout: 20000 });
await page.pdf({ path: 'output.pdf' });

Replace main article with a selector that exists on the target page. If the content appears after a known delay and there is no useful selector, a bounded delay can be used, but it is less robust than a page-specific condition. The PDF method waits for fonts to load by default; if text still renders incorrectly, inspect font loading in the page and the output separately from navigation timing.

4. Control print layout and page appearance

Puppeteer’s PDF output uses the page’s print CSS media type by default. That means print-specific rules can hide navigation, change colors, or rearrange content. If you need screen styles instead, emulate the screen media type before creating the PDF:

await page.emulateMediaType('screen');
await page.pdf({ path: 'output.pdf' });

Print CSS is often the better choice for documents intended to be printed. If a color disappears or changes, inspect the site’s print styles. Puppeteer notes that browsers may modify colors for printing; the CSS property -webkit-print-color-adjust can request exact color rendering.

Set PDF options explicitly when output consistency matters. For example, this writes landscape A4, includes background graphics, and sets margins:

await page.pdf({
  path: 'output.pdf',
  format: 'A4',
  landscape: true,
  printBackground: true,
  margin: {
    top: '12mm',
    right: '12mm',
    bottom: '12mm',
    left: '12mm',
  },
});

Common Puppeteer PDF options include:

Option Effect Use it when
format Chooses a named paper size such as A4 or Letter You need a predictable page size
landscape Rotates the page orientation Wide tables or diagrams need more width
margin Sets top, right, bottom, and left margins Content needs room around the page edge
pageRanges Limits output to selected page numbers or ranges You need only part of a long document
scale Scales the printed page content You need to fit or enlarge content, with layout tradeoffs
printBackground Includes background graphics Colors or background images matter to the result
preferCSSPageSize Lets CSS @page size take priority The page defines its own paper dimensions
path Writes the PDF to a file You want a local artifact

For option details and accepted values, use the Puppeteer PDFOptions reference. The tagged option is experimental in the documented API; do not assume it produces accessible or correctly tagged output without checking the generated document.

5. Or skip the browser setup

ScreenshotNeo is a website screenshot API that can return a PDF from one GET request. The endpoint is https://api.screenshotneo.com/v1/shot. Here is a cURL request using the documented API pattern:

A clean capture can remove common overlays before the page is rendered.
A clean capture can remove common overlays before the page is rendered.
curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

See the ScreenshotNeo API documentation for authentication and PDF parameters. The request uses access_key and a target url; use the documented PDF format option for the returned file.

The same endpoint can be called from Python or Node.js. These examples follow ScreenshotNeo’s documented request shape and save the response body to a file:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://stripe.com",
        "format": "pdf",
    },
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await require('node:fs/promises').writeFile('page.pdf', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response includes X-Page-Verdict and X-Billed headers. It also provides an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools.

The free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; the other monthly tiers are Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.

6. Troubleshooting PDF generation

Symptom Likely cause Fix
No PDF appears Chrome cannot write to the current directory, or the command is not using the expected executable Check the working directory’s write permissions and run Chrome by its full path if needed.
Headers or footers are present Chrome adds print metadata by default Use --no-pdf-header-footer; on older versions, check whether --print-to-pdf-no-header applies.
Content is missing The page had not rendered the content before printing, or content depends on interaction Increase a bounded timeout, wait for a meaningful selector in Puppeteer, or reproduce the required interaction before generating the PDF.
Output is blank Navigation failed, the page is blank to the browser, or the site requires authentication or JavaScript state Check the page in a normal browser, inspect navigation errors, and provide the needed session or page setup in your automation.
Images are missing Images are lazy-loaded or resource requests failed Wait for the relevant content, check image requests, and scroll or otherwise trigger lazy loading where the page requires it.
Fonts or text look wrong A web font failed to load or a print rule changes typography Check font requests and print CSS. Puppeteer waits for fonts during PDF generation by default, but that cannot make an unavailable font load.
Colors or backgrounds differ Print media styles or print color adjustment changed the appearance Set printBackground: true when needed, inspect print CSS, and use -webkit-print-color-adjust where exact colors are required.
Navigation times out The chosen wait condition never becomes satisfied, often because requests remain active Choose a more suitable navigation condition and then wait for the page-specific element that indicates readiness.
PDF pages break awkwardly Print CSS, paper size, margins, or scaling do not suit the content Adjust @media print or @page rules and test the paper format, margins, and scale.

7. Performance, reliability, and cost

A local headless browser avoids a per-request hosted conversion charge, but the machine running it must have Chrome installed, enough resources for the pages being rendered, and permission to write the output. Puppeteer adds browser lifecycle and dependency management. For repeated jobs, close browsers in a finally block, set navigation and selector timeouts, and record the target URL and failure reason so a bad conversion can be diagnosed.

Capture time depends on the page and its resources; the cited documentation does not provide a universal conversion time or guarantee for arbitrary URLs. Waiting longer may allow slow content to arrive, but can reduce throughput. Waiting for an overly broad network-idle condition can stall on persistent requests, while printing too early can omit content. Pick a readiness rule based on the target site and validate representative PDFs.

With a hosted API, consider request volume, plan limits, response handling, and retries. ScreenshotNeo’s free tier is 1,000 shots per month; paid plans start at $5 for 3,000, with larger plans listed above. Its billing headers let a caller distinguish outcomes, including unbilled failures and cache hits. Treat timeouts as uncertain outcomes: retry selectively, and avoid uncontrolled parallel requests. See the service documentation for current parameters and usage details.

8. A practical checklist

  1. Choose manual browser printing, Chrome Headless, Puppeteer, or a hosted API based on how often the job runs.
  2. Confirm the page is publicly reachable, or provide the authentication and cookies required to render it.
  3. Decide whether the PDF should follow print styles or screen styles.
  4. Set paper size, orientation, margins, backgrounds, and page range as needed.
  5. Wait for a meaningful readiness condition; do not assume one generic timeout covers every site.
  6. Check the generated file for missing content, font issues, unexpected pagination, and unwanted headers or footers.
  7. For a recurring workflow, handle browser cleanup or API errors and retain enough logs to diagnose failures.

9. Frequently asked questions

Can I generate a PDF from a URL without installing software?

Use the print-to-PDF option in a browser you already have, or call a hosted URL-to-PDF API. The browser route is simplest for occasional use; code or an API is more repeatable.

Will the PDF look exactly like the page on screen?

Not necessarily. Puppeteer uses print media by default, and print CSS can change layout and colors. Emulate screen media when that is the intended output, then inspect the PDF.

Can I save a page that requires login?

Only if the browser session used for capture has access. In automation, configure the required authentication or session state; a bare public URL does not include your login.

Is a generated PDF guaranteed to be accessible?

No. PDF output and experimental tagged-PDF controls do not by themselves establish that a document is accessible. Check the resulting file with appropriate accessibility tools.

Sources