ScreenshotNeo

BlogHTML to image & PDF

How to Convert a Web Page to PDF in Bash

Convert any web page to PDF from Bash with headless Chrome, WeasyPrint, or wkhtmltopdf, including timing, styling, security, and automation tips.

By the ScreenshotNeo team1 October 20267 min read

The simplest reliable command is headless Chrome:

chrome --headless --print-to-pdf https://example.com

Chrome writes output.pdf in the current directory. To choose the filename and remove the printed date, URL, and page number, use:

chrome --headless --print-to-pdf=page.pdf --no-pdf-header-footer https://example.com

Headless Chrome is usually the best starting point for modern sites because it runs the same browser rendering model used by Chrome, including client-side JavaScript. The command-line options below come from Chrome’s headless documentation: Chrome Headless mode.

1. Convert a page with Chrome or Chromium

Basic conversion

# Chrome
chrome --headless --print-to-pdf=page.pdf https://example.com

# Chromium on many Linux distributions
chromium --headless --print-to-pdf=page.pdf https://example.com

The executable may be named google-chrome, google-chrome-stable, or chromium depending on your operating system. Check which one is installed:

command -v google-chrome || command -v google-chrome-stable || command -v chromium || command -v chromium-browser

A reusable Bash script

#!/usr/bin/env bash
set -Eeuo pipefail

url=${1:?Usage: $0 URL [OUTPUT.pdf]}
out=${2:-output.pdf}

browser=$(command -v google-chrome \
       || command -v google-chrome-stable \
       || command -v chromium \
       || command -v chromium-browser \
       || true)

if [[ -z "$browser" ]]; then
  echo "No Chrome or Chromium executable found" >&2
  exit 1
fi

"$browser" \
  --headless \
  --disable-gpu \
  --no-sandbox \
  --print-to-pdf="$out" \
  --no-pdf-header-footer \
  "$url"

if [[ ! -s "$out" ]]; then
  echo "PDF was not created or is empty: $out" >&2
  exit 1
fi

printf 'Wrote %s (%s bytes)\n' "$out" "$(wc -c < "$out")"

Use --no-sandbox only in an isolated environment where you understand the security trade-off. Do not add it automatically to a multi-tenant or untrusted workload.

Pages that need more time

Some pages render an initial shell and fill it with JavaScript later. Chrome documents --timeout=5000 as a maximum capture wait of five seconds:

chrome --headless \
  --timeout=5000 \
  --print-to-pdf=page.pdf \
  https://example.com

For timer-driven content, --virtual-time-budget=42000 advances JavaScript timers in virtual time:

chrome --headless \
  --virtual-time-budget=42000 \
  --print-to-pdf=page.pdf \
  https://example.com

These flags control waiting; they do not guarantee that every network request, login flow, animation, or application-specific readiness condition has completed. For a critical workflow, inspect the resulting PDF and add a page-specific readiness strategy.

Useful Chrome options

Need Option Example
Choose output --print-to-pdf=FILE --print-to-pdf=invoice.pdf
Remove print metadata --no-pdf-header-footer Suppresses date, URL, and page number
Wait up to a duration --timeout=MS --timeout=5000
Advance timers --virtual-time-budget=MS --virtual-time-budget=42000

Older Chrome installations may use the historical spelling --print-to-pdf-no-header for header suppression. If the current spelling is rejected, check the installed version’s help output.

2. Convert with WeasyPrint

WeasyPrint’s command-line interface accepts a URL and an output filename:

weasyprint https://example.com page.pdf

It is a good fit when you need HTML and CSS document layout rather than full browser fidelity. You can add a stylesheet from Bash with process substitution:

weasyprint https://example.com page.pdf \
  -s <(echo 'body { font-family: serif !important }')

WeasyPrint may report unsupported CSS properties on standard error. Review those warnings when layout matters. Its documentation also warns that untrusted HTML or CSS can create security problems. Treat remote HTML, CSS, images, and fonts as untrusted input and isolate the conversion process.

3. Convert with wkhtmltopdf

The basic command is:

wkhtmltopdf https://example.com page.pdf

The wkhtmltopdf manual documents page size, margins, orientation, JavaScript settings, JavaScript delay, print media, headers and footers, and multiple page objects:

wkhtmltopdf \
  --page-size A4 \
  --orientation Portrait \
  --margin-top 15mm \
  --margin-right 15mm \
  --margin-bottom 15mm \
  --margin-left 15mm \
  --print-media-type \
  --javascript-delay 2000 \
  https://example.com page.pdf

wkhtmltopdf is a legacy option to evaluate carefully. The official downloads page identifies 0.12.6 as the stable series, released June 11, 2020, and warns against processing untrusted HTML or JavaScript because it can compromise the server. Check your distribution’s package and build behavior before standardizing on it.

4. Choosing the renderer

Renderer Choose it when Watch for
Chrome/Chromium The page uses modern JavaScript, web fonts, responsive layout, or client-side rendering. Readiness, authentication, and network completion still require verification.
WeasyPrint You want direct HTML/CSS-to-PDF conversion and stylesheet overrides. Some CSS properties are unsupported; it is not pixel-identical browser rendering.
wkhtmltopdf You need its documented paper, margin, header/footer, and JavaScript-delay controls in an existing legacy workflow. Its stable series is dated; review security and package behavior.

5. Make the output predictable in automation

Use safe filenames

slug=$(printf '%s' "$url" | sed -E 's#[^A-Za-z0-9._-]+#_#g')
out="${slug:-page}.pdf"
chrome --headless --print-to-pdf="$out" "$url"

Keep the URL quoted. Never concatenate untrusted URL text into shell syntax.

Check that a PDF was produced

if [[ ! -s page.pdf ]]; then
  echo "missing or empty PDF" &2
  exit 1
fi

# Optional file-type check when file(1) is available
file page.pdf

Process a list of URLs

#!/usr/bin/env bash
set -Eeuo pipefail

while IFS= read -r url; do
  [[ -z "$url" ]] && continue
  name=$(printf '%s' "$url" | sed -E 's#^[^:]+://##; s#[^A-Za-z0-9._-]+#_#g')
  chrome --headless --print-to-pdf="${name}.pdf" "$url"
done < urls.txt

Run independent browser processes with a controlled concurrency limit rather than starting an unlimited number. Browser startup, memory, fonts, and network bandwidth become bottlenecks quickly.

6. Authentication, private pages, and untrusted input

A public URL is the easiest case. Private pages may require cookies, HTTP authentication, VPN access, or an application login flow. Chrome’s simple print command does not provide a universal, safe way to inject every kind of session state. If authentication is required, use a controlled browser profile or an application-specific automation layer, and keep credentials out of command history and logs.

Do not feed arbitrary user-supplied URLs to a privileged renderer without isolation. A renderer may request internal network addresses, access local files, execute JavaScript, or consume excessive CPU and memory. Use an allowlist where possible, run as a low-privilege user, restrict egress, apply timeouts, and clean temporary profiles.

7. Troubleshooting

Symptom Likely cause Fix
command not found Chrome, Chromium, WeasyPrint, or wkhtmltopdf is not installed or is not on PATH. Install the selected renderer or pass its absolute executable path.
PDF is blank The page is client-rendered, blocked, or captured before content appears. Try Chrome, increase --timeout, use a virtual-time budget, and verify the URL in a browser.
Images are missing Images loaded after capture, require authentication, or fail due to network/CORS policy. Allow more time, verify image URLs, and ensure the renderer can reach the assets.
Fonts differ The required web font is unavailable or still loading. Install the font in the rendering environment, wait longer, or use a fallback.
Headers and footers appear Chrome print metadata is enabled. Use --no-pdf-header-footer; older versions may require --print-to-pdf-no-header.
Layout is cut off Print CSS, viewport assumptions, fixed elements, or page-break rules differ from screen layout. Inspect print styles, choose the appropriate renderer, and review page size and margins.
Command hangs A request, script, or page load never completes. Set a timeout, isolate the process, and capture diagnostics from stderr.
Permission or sandbox error The browser cannot create its profile or sandbox in the runtime. Set a writable temporary directory and correct user permissions. Use --no-sandbox only in an appropriately isolated environment.
WeasyPrint CSS warnings The stylesheet uses unsupported properties. Read stderr and provide a compatible stylesheet or use a browser renderer.

8. Performance, reliability, and cost

  • Startup cost: launching a browser per URL is simple but expensive. For batches, a long-lived browser controlled by an automation framework can reduce startup overhead.
  • Resource limits: set CPU, memory, process, and network limits. Large pages, long screenshots, video, and infinite-scroll scripts can exhaust a worker.
  • Retries: retry transient network failures with a small exponential backoff. Do not blindly retry deterministic errors such as an invalid URL or a blocked domain.
  • Reproducibility: pin the browser or renderer version, fonts, locale, timezone, and CSS inputs when PDFs are used as build artifacts.
  • Verification: check file existence and size, then inspect representative PDFs for pagination, images, fonts, and final content. A zero exit status alone does not prove the page rendered correctly.
  • Cost: self-hosted tools shift cost to your compute, storage, bandwidth, maintenance, and security operations. A hosted API shifts browser operations to the provider and charges according to its plan and billing rules.

Or skip the browser setup

ScreenshotNeo provides a website capture API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. The API accepts the URL and capture options, while the service handles browser setup and page preparation. See the ScreenshotNeo API documentation for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = await res.arrayBuffer();
await Bun.write('shot.webp', data);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and whether it was billed. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. You can use full-page capture, CSS-element capture, custom CSS and JavaScript, waiting rules, headers and cookies, PDF paper settings, caching, signed links, async webhooks, bulk capture, and other options on every plan.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.

9. FAQ

Does Chrome save the PDF in the current directory?

Yes. Without an explicit filename, Chrome documents output.pdf in the current working directory.

Can I convert a local HTML file?

Use a file URL such as file:///absolute/path/page.html, subject to the renderer’s local-file and resource permissions. Treat local HTML and referenced assets as trusted only when you control them.

Which tool renders JavaScript best?

Headless Chrome or Chromium is the appropriate first choice for modern client-rendered pages. You still need to verify readiness for each application.

Why does the screen view differ from the PDF?

PDF generation uses print layout. Print media rules, page size, margins, fixed elements, and page-break rules can change the result.

How do I remove Chrome’s URL and date?

Pass --no-pdf-header-footer. On older installations, try --print-to-pdf-no-header.