ScreenshotNeo

BlogHow-to

Convert a Website URL to PDF From the Command Line

Print a website URL to PDF with Chrome’s headless CLI, then use Puppeteer when you need repeatable jobs or control over PDF layout.

By the ScreenshotNeo team4 October 20266 min read

To convert a website URL to PDF from the command line, use Chrome’s headless print option:

chrome --headless --print-to-pdf https://example.com/

Chrome writes output.pdf to the current working directory by default. To suppress the printed URL, date, and page headers and footers, add --no-pdf-header-footer. These headers are separate from the page’s own print styling. See Chrome’s headless documentation.

This is the simplest method for a one-off capture. Use Puppeteer when the job needs to run repeatedly, wait for a specific page condition, or set PDF options such as paper size and margins.

1. Print a URL with Chrome’s command-line interface

Run the command in a terminal where Chrome is installed:

chrome --headless --print-to-pdf https://example.com/

The PDF is saved as output.pdf in the current directory. Use an explicit output path when you want a different filename or location:

chrome --headless --print-to-pdf=/tmp/example.pdf https://example.com/

To omit Chrome’s generated headers and footers:

chrome --headless --print-to-pdf --no-pdf-header-footer https://example.com/

Some older Chrome versions use the earlier flag spelling --print-to-pdf-no-header. If Chrome rejects the newer option, check the installed version and its supported flags.

Wait for delayed page content

For a page that populates after navigation, Chrome’s --timeout option waits up to the specified number of milliseconds before capture. For example:

chrome --headless --print-to-pdf --timeout=5000 https://example.com/

This is a maximum wait, not proof that the page is ready. A fixed delay may not be enough for scripts, images, fonts, or access checks.

For pages that rely on timers, Chrome documents a virtual time budget:

chrome --headless --print-to-pdf --virtual-time-budget=42000 https://example.com/

Treat virtual time as a rendering aid. Review the resulting PDF for content that appeared too late or did not load.

2. Use Puppeteer for repeatable PDF generation

Puppeteer is useful when you need a script, explicit PDF settings, or a readiness condition before printing. Install it in a Node.js project:

npm install puppeteer

Save this as save-pdf.mjs. It accepts the URL and output path as command-line arguments:

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com/';
const output = process.argv[3] ?? 'output.pdf';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle2', timeout: 60_000 });
  await page.pdf({
    path: output,
    format: 'A4',
    printBackground: true,
    displayHeaderFooter: false,
    margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
  });
} finally {
  await browser.close();
}

Run it with:

node save-pdf.mjs https://example.com/ ./example.pdf

Puppeteer’s page.pdf() returns PDF bytes and uses the print CSS media type by default. If the site’s screen layout is more appropriate, set screen media before calling pdf():

await page.emulateMediaType('screen');
await page.pdf({ path: output, printBackground: true });

For exact printed colors, Puppeteer documents the CSS property -webkit-print-color-adjust. Page-authored print styles can change layout or hide content, so check the page’s CSS if the PDF differs from the browser view. See Puppeteer’s PDF API.

Choose a readiness condition that fits the page

networkidle2 waits for network activity to settle, but pages with polling, analytics, or long-lived requests may not reach that state. If the page exposes a meaningful element once its main content is ready, wait for that selector instead:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.waitForSelector('main article', { timeout: 30_000 });
await page.pdf({ path: output, format: 'A4', printBackground: true });

Replace main article with a selector that represents the actual content on the target page. A selector wait is only useful if that element appears after the content you need is ready.

3. Pick the right workflow

Need Use
One URL, one PDF, minimal setup Chrome CLI with --headless --print-to-pdf
Repeated captures or configurable margins and paper size Puppeteer and page.pdf()
Existing Playwright project Use its browser and page workflow, and check the selected browser’s current PDF behavior

Playwright’s documentation covers page navigation and notes that headless mode does not support navigation to a PDF document. That caveat is about opening a PDF URL; it is not a blanket statement about every way to generate a PDF. See Playwright’s Page API.

4. Layout, output, and edge cases

  • Print CSS: Browser PDF generation applies print behavior. A site may hide menus, change colors, or rearrange sections in print styles.
  • Backgrounds and colors: In Puppeteer, enable printBackground if you need background graphics. Page CSS can also affect color output.
  • Page breaks: Long pages can split content awkwardly. Inspect page boundaries and adjust the site’s print CSS where you control it.
  • Dynamic content: Lazy images or script-rendered sections may be missing if capture happens before they appear. Use an appropriate wait and inspect the output.
  • Authentication and access checks: A command-line browser may receive a sign-in page, challenge, or error page instead of the expected content. The resulting PDF reflects what the browser could access.
  • URL handling: Quote URLs in your shell when they contain characters interpreted by the shell, such as ampersands.

For example, quote a URL with query parameters:

chrome --headless --print-to-pdf 'https://example.com/report?month=may&view=full'

5. Troubleshooting

Symptom Likely cause What to do
chrome: command not found Chrome is not on the shell’s PATH, or the executable has a platform-specific name. Locate the installed Chrome executable and invoke it by its full path.
Unknown or rejected PDF flag The installed version uses different flag support or an older spelling. Check the Chrome version and use the flag name documented for that binary. Older versions may need --print-to-pdf-no-header.
PDF contains a blank or incomplete page The page had not rendered its content when capture occurred, or access checks prevented loading. Open the URL in a browser to confirm what is accessible; then add an appropriate wait or use a script that waits for a content selector.
Colors or layout differ from the screen Print media CSS changes presentation; backgrounds may not be included. Inspect print preview and print styles. With Puppeteer, try screen media or enable printBackground as needed.
Command runs but PDF is not where expected Chrome wrote the default output.pdf in the current working directory. Check that directory or specify an output file path with the print-to-PDF option.
Script hangs waiting for network idle The site keeps requests open or generates continuous traffic. Use a selector or another page-specific readiness condition, with a finite timeout.

6. Version and reliability notes

Chrome’s headless documentation says headless and headful modes now share Chrome’s implementation. Since Chrome 132.0.6793.0, the old Headless mode is available only as the separate chrome-headless-shell binary. Old tutorials may assume that mode is included in the regular Chrome binary; check which executable and version your environment provides. See Chrome’s headless documentation.

For automation, use finite navigation and readiness timeouts, close the browser in a finally block, and inspect generated PDFs when page content matters. Browser printing does not guarantee identical output across sites, operating systems, or browser versions.

7. Or skip the browser setup

If you need a PDF from an API call, ScreenshotNeo accepts a URL and returns a PDF. The API supports PDF settings such as paper size, margins, landscape orientation, and page ranges. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', bytes));

ScreenshotNeo can accept cookie and consent banners before capture and remove more than 60 known consent platforms, along with newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture.

The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Learn more at ScreenshotNeo or sign up for the free plan.

FAQ

Does Chrome save the PDF as the page title?

By default, Chrome writes output.pdf in the current directory. Set an output path in the print-to-PDF option to choose another name.

Can I make the PDF show the screen version of a page?

With Puppeteer, call page.emulateMediaType('screen') before page.pdf(). The default PDF rendering uses print media.

Does a longer timeout guarantee a complete capture?

No. A delay gives content more time to appear, but it cannot guarantee that scripts, images, or access checks finish. Use a relevant readiness condition and review the PDF.