ScreenshotNeo

BlogHow-to

Using a CLI to Generate PDFs Locally

Generate PDFs locally with Pandoc, Chrome Headless, or WeasyPrint. Choose the right renderer, handle dependencies, automate safely, and troubleshoot failures.

By the ScreenshotNeo team1 October 20267 min read

Use Pandoc when your input is Markdown or another source document, Chrome Headless when you need to print an existing web page, and WeasyPrint when you want an HTML-to-PDF command-line workflow. The renderer you choose determines CSS support, JavaScript behavior, fonts, dependencies, and security considerations.

Choose the right local PDF command

Input Start with Why
Markdown, DOCX, or another document format Pandoc Converts source documents and delegates PDF creation to a selectable PDF engine.
A live or existing web page Chrome Headless Prints the page with a browser renderer, including browser layout and JavaScript.
HTML with print CSS WeasyPrint Provides an HTML-to-PDF CLI route without requiring a full interactive browser.

There is no universal winner. Decide based on the input format, whether JavaScript must run, how much CSS control you need, font availability, PDF standards, and whether the content is trusted.

1. Generate a PDF from Markdown with Pandoc

The shortest command is:

pandoc input.md -o output.pdf

Pandoc normally uses LaTeX for PDF output, so a LaTeX engine must be installed. See the Pandoc User’s Guide and installation documentation for current requirements.

Select a different PDF engine

pandoc input.md -o output.pdf --pdf-engine=PROGRAM

The Pandoc manual documents engines and output routes including weasyprint, wkhtmltopdf, pagedjs-cli, and prince. The engine affects layout features, dependencies, and security behavior.

Use a table of contents and metadata

pandoc input.md -o output.pdf --toc --metadata title='Project guide' --metadata author='Example Team'

Apply a CSS stylesheet through HTML output

For HTML-oriented styling, create an intermediate HTML file and then print or convert it with an HTML-capable engine:

pandoc input.md -s --css=print.css -o intermediate.html
pandoc intermediate.html -o output.pdf --pdf-engine=weasyprint

Verify the selected engine’s CSS and pagination support before relying on advanced rules such as page breaks, running headers, or generated content.

2. Print a web page with Chrome Headless

Chrome’s documented headless printing pattern is:

chrome --headless --print-to-pdf https://example.com/

Chrome writes output.pdf in the current directory. Add --no-pdf-header-footer when you do not want the browser’s print header and footer.

chrome --headless --no-pdf-header-footer --print-to-pdf=page.pdf https://example.com/

Allow delayed content to render

Use the documented timeout and virtual-time options when the page needs time to load or run timers:

chrome --headless --timeout=30000 --virtual-time-budget=5000 --print-to-pdf=page.pdf https://example.com/

These flags set limits or advance page timers; they do not guarantee that every application has finished rendering. For authenticated pages, provide the browser profile or session setup required by your environment and keep credentials out of shell history.

3. Generate a PDF from HTML with WeasyPrint

WeasyPrint provides a command-line HTML-to-PDF path. A typical invocation is:

weasyprint input.html output.pdf

Consult the current WeasyPrint first-steps documentation for installation requirements and CLI options. Treat this as an HTML renderer with its own CSS support; it does not provide the same JavaScript execution model as a browser.

4. Complete runnable examples

Shell script: choose Pandoc, Chrome, or WeasyPrint

#!/usr/bin/env sh
set -eu

input=${1:?usage: ./make-pdf.sh INPUT}
output=${2:-output.pdf}

case "$input" in
  *.md|*.markdown)
    pandoc "$input" -o "$output"
    ;;
  *.html|*.htm)
    weasyprint "$input" "$output"
    ;;
  http://*|https://*)
    chrome --headless --no-pdf-header-footer --print-to-pdf="$output" "$input"
    ;;
  *)
    echo 'Use Markdown, HTML, or an HTTP(S) URL.' >&2
    exit 2
    ;;
esac

Python: run a local renderer

from pathlib import Path
import subprocess

source = Path('input.md')
output = Path('output.pdf')
subprocess.run(['pandoc', str(source), '-o', str(output)], check=True)
print(output.resolve())

Node.js: print a URL with Chrome

import { spawn } from 'node:child_process';

const output = 'page.pdf';
const url = 'https://example.com/';
const child = spawn('chrome', [
  '--headless',
  '--no-pdf-header-footer',
  `--print-to-pdf=${output}`,
  url
], { stdio: 'inherit' });
child.on('exit', code => process.exit(code ?? 1));

5. Dependencies and installation checks

  1. Install the primary tool using its official installation instructions.
  2. Check the executable is available: pandoc --version, chrome --version, or weasyprint --version.
  3. For Pandoc PDF output, install and verify the selected PDF engine.
  4. Install the fonts your document requires and confirm they are visible to the renderer.
  5. Run a small sample before processing a large batch.

Pandoc itself does not remove the PDF-engine dependency. Package names and commands vary by operating system, so use the current upstream installation pages rather than copying an old package command.

6. Layout, fonts, and page-break control

  • Use print-specific CSS for HTML workflows, including @page, margins, and explicit page-break rules supported by your renderer.
  • Embed or install required fonts; a missing font can change line wrapping and page count.
  • Keep images at appropriate resolution and use stable absolute paths or URLs.
  • Validate links, headings, tables, and images in the generated PDF instead of assuming source HTML guarantees the output.
  • If PDF/A, PDF/UA, tagging, or accessibility matters, verify the exact renderer version and validate the generated file. A command-line flag alone is not proof of compliance.

7. Security when converting untrusted input

PDF engines can access external resources and may expose local data if given unsafe options. Pandoc’s documentation specifically discusses risks around engines such as wkhtmltopdf, including local-file exposure through file: URIs and SSRF scenarios involving raw HTML. Treat input documents, templates, renderer options, and network resources as separate trust boundaries.

  • Do not process untrusted files with a powerful renderer on a machine that contains secrets.
  • Run conversions in a restricted account or container with limited filesystem and network access.
  • Disable or restrict external resource loading where your renderer supports it.
  • Audit engine options before accepting user-controlled arguments.
  • Never pass unsanitized user input directly into a shell command; use argument arrays in Python or Node.js.

8. Troubleshooting

“pdflatex not found” or a similar Pandoc engine error

Cause: Pandoc is installed but its default LaTeX engine is missing. Fix: install a LaTeX distribution or select an installed engine with --pdf-engine=PROGRAM.

Chrome creates an empty or incomplete PDF

Cause: the page has delayed JavaScript, blocked resources, authentication requirements, or a loading error. Fix: try --timeout and --virtual-time-budget, confirm the URL works in the same environment, and provide the required session setup.

Headers or footers appear unexpectedly

Cause: Chrome’s print decoration is enabled. Fix: add --no-pdf-header-footer.

Fonts or line breaks differ between machines

Cause: different installed fonts, font versions, or fallback behavior. Fix: install or bundle the intended fonts and run the same renderer version in a controlled environment.

Images are missing

Cause: relative paths, blocked network requests, unsupported formats, or permissions. Fix: use valid paths or URLs, verify access from the conversion process, and inspect renderer logs.

CSS works in a browser but not in the PDF

Cause: the selected engine supports a different subset of CSS or does not execute the page’s JavaScript. Fix: use Chrome for browser-dependent pages, or simplify and test print CSS for WeasyPrint or another Pandoc engine.

The command hangs

Cause: a network resource, script, or renderer process never finishes. Fix: set a bounded timeout where available, isolate external resources, and terminate stuck worker processes in your automation.

9. Performance, reliability, and cost

These tools run locally, so direct service charges are generally replaced by your own CPU, memory, storage, dependency maintenance, and operations. The documentation does not establish a universal speed or resource winner.

  • Reuse a warm browser process only if your isolation model permits it; otherwise launch clean processes for stronger separation.
  • Cache stable source files, fonts, and remote assets to reduce repeated network work.
  • Use bounded timeouts and retries for network-backed pages, but avoid retrying deterministic syntax or missing-dependency errors.
  • Record renderer versions, command arguments, exit codes, and output hashes for reproducibility.
  • For batches, limit concurrency to the memory available to your browser or PDF engine.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API that can also return PDFs. Use the same one-call workflow from cURL, Python, or Node.js, and see the API documentation for PDF options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o page.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("page.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. It also offers an MCP server for Claude, Cursor, and other MCP clients, plus full-page capture, PDF settings, custom CSS and JavaScript, waiting controls, blocking rules, authentication options, caching, signed links, async jobs, bulk capture, usage data, and an OpenAPI specification. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.

FAQ

Can Pandoc convert a web page directly?

It can process supported source formats, but a live JavaScript-heavy page is usually better handled by a browser printer such as Chrome Headless.

Which tool should run in CI?

Use the renderer that matches your input and pin its version, fonts, and dependencies. Add output validation so layout changes fail visibly.

Can I guarantee PDF accessibility from one command?

No. Accessibility and PDF standards depend on the renderer, version, source structure, and validation. Generate the file, then validate it against the required standard.

Is a local CLI better than an API?

A local CLI gives direct control over dependencies and data locality. An API avoids browser and renderer setup and can provide managed capture behavior; choose according to your operational constraints.