ScreenshotNeo

BlogHTML to image & PDF

How to Archive a Webpage as a PDF with wkhtmltopdf

Convert a webpage to PDF with wkhtmltopdf, choose page layout options, and troubleshoot dynamic content, authentication, and missing resources.

By the ScreenshotNeo team4 October 20268 min read

To archive one webpage as a PDF with wkhtmltopdf, install the command-line tool and run wkhtmltopdf 'https://example.com/page' 'page.pdf'. The first argument is the page URL; the second is the output filename. Then open the PDF and check its text, images, layout, links, and page breaks. Results depend on the page and the wkhtmltopdf build installed on your system.

1. Install wkhtmltopdf and check your build

The project provides a precompiled binary or the option to build from source. Choose an installation route that fits your operating system, then check the installed version and available options:

wkhtmltopdf --version
wkhtmltopdf --help

Builds and packaged manuals can differ. Use the manual associated with your installation when an option behaves differently from an example here. The project documentation describes the command as a headless tool that renders HTML into PDF using Qt WebKit. Project overview · Command-line manual

2. Convert a webpage to PDF

Run the basic command, replacing the example URL and filename with your target and desired output path:

wkhtmltopdf 'https://example.com/page' 'page.pdf'

The documented command form is wkhtmltopdf [GLOBAL OPTION]... [OBJECT]... <output file>. For a single page, the URL and output filename are enough. Quote paths and URLs so the shell treats spaces and punctuation as part of the argument.

When the command finishes, open the PDF. Check that the expected content loaded and that page breaks, image placement, links, and page dimensions work for your purpose. Conversion success only means the command produced an output file; it does not establish that every remote resource or dynamic element appeared.

3. Set paper size, orientation, and margins

The documented defaults are A4 paper and portrait orientation. Set layout options when the PDF must fit a particular paper format or a wide page:

wkhtmltopdf \
  --page-size Letter \
  --orientation Landscape \
  --margin-top 15mm \
  --margin-bottom 18mm \
  --margin-left 12mm \
  --margin-right 12mm \
  'https://example.com/page' \
  'page.pdf'
Need Option Notes
Choose paper --page-size A4, A3, Letter, or Legal A4 is the documented default. Pick a size for the intended use.
Set custom dimensions --page-width and --page-height Use when a named paper size does not fit the page.
Use landscape --orientation Landscape Useful to try for wide tables or layouts; inspect the resulting PDF.
Adjust whitespace --margin-top, --margin-bottom, --margin-left, --margin-right Set enough room for content and any headers or footers.

For precise sizing, consult the installed build’s help or manual for accepted dimensions and units. Changing paper size or margins can shift content and page breaks, so review the output after each layout change.

4. Choose screen or print styles and backgrounds

By default, the documented behavior selects screen media styles. Add --print-media-type to request print styles instead:

wkhtmltopdf --print-media-type 'https://example.com/page' 'print-styled.pdf'

This changes which CSS media rules are selected; it cannot guarantee that a site’s print stylesheet produces a complete or clean document. Compare the output with and without the option if the page has different screen and print layouts.

Background printing is enabled by default in the documented manual. Use --no-background to omit background graphics:

wkhtmltopdf --no-background 'https://example.com/page' 'no-background.pdf'

Print media selection and background printing are separate choices. Check the PDF for missing color blocks, background images, or text that depends on the selected stylesheet.

5. Add headers and footers

Headers and footers can use text or HTML. For a simple page count, leave enough top and bottom margin and add a footer:

wkhtmltopdf \
  --margin-bottom 20mm \
  --footer-right 'Page [page] of [topage]' \
  'https://example.com/page' \
  'page-numbered.pdf'

Documented substitution variables include [page], [topage], [webpage], [title], and [isodate]. The command-line manual also describes HTML header and footer options. Confirm exact option names and behavior in the manual for your installed build.

6. Handle JavaScript-rendered pages

A page may fill in its content after the initial HTML loads. The CLI documents JavaScript controls, a configurable JavaScript delay, a wait-until-window-status option, and a script hook that runs after page load. These controls can help when content appears late, but they do not ensure compatibility with every dynamic site.

For example, try a delay when the page needs extra time to populate:

wkhtmltopdf --javascript-delay 2000 'https://example.com/page' 'delayed.pdf'

The delay value is in milliseconds. Choose a value that fits the page rather than assuming one wait time works for all URLs. If the page exposes a suitable window status, the documented --window-status option can wait for that status. A post-load script can also be configured with the manual’s JavaScript options. Consult your build’s help for the precise syntax.

  1. First capture with the basic command and inspect the PDF.
  2. If expected content is absent, determine whether it appears only after scripts or a delayed request run in a normal browser.
  3. Try an appropriate delay or documented page-status/script option, then inspect the new output.
  4. If the content still does not render as needed, try a browser’s built-in print-to-PDF path and compare the result.

7. Convert pages that require access or load remote resources

The manual documents HTTP authentication, custom headers, and proxy settings. These are configuration controls, not a guarantee that a protected page will be accessible. Use them only for pages you are authorized to access. Avoid placing secrets directly in shell history or shared scripts; use your environment’s secret-handling practices.

Remote images, fonts, stylesheets, scripts, and other resources can affect the PDF. If an asset is missing, check that the converter can reach its URL and that access requirements are satisfied. The manual documents load-error and media-error handling options; consult the installed manual to decide whether to ignore or stop on particular resource errors. Ignoring an error can produce an incomplete document.

For local HTML input, wkhtmltopdf has local-file access controls, including an option to disable access with an allow-list mechanism and an option to enable it. Keep local access as narrow as the task permits. Do not grant broad access to local files without a specific reason.

8. Combine multiple documents (optional)

The command accepts page, cover, and table-of-contents objects in sequence. This supports compiling several inputs into one PDF, but is unnecessary for a single webpage. A cover omits headers and footers and is excluded from the table of contents; the table of contents is generated using XSLT. See the official command-line manual for object syntax and related options.

9. Troubleshooting

Symptom Likely cause What to try
No PDF appears The output path is invalid, unwritable, or its parent directory does not exist. Choose a writable output path, create the directory, and check the command’s error output.
PDF exists but expected content is blank or missing The page may depend on JavaScript, late-loading content, authentication, or remote resources. Check access in a browser, then try a documented delay, window-status wait, authentication/header setting, or browser print-to-PDF fallback as appropriate.
Images, fonts, or styles are missing A resource URL failed, requires access, or is blocked by network/proxy configuration. Check resource reachability and authorized access; review load-error and proxy settings in the installed manual.
Layout looks wrong compared with the browser Paper size, orientation, margins, or selected screen/print CSS differ from what the page needs. Try an appropriate paper size and orientation, adjust margins, and compare screen output with --print-media-type.
Backgrounds are absent The command may have disabled backgrounds, or the chosen styles do not define the expected background. Remove --no-background if present and check both the selected media type and the PDF.
Header or footer overlaps content There is not enough margin reserved for it. Increase the corresponding margin and inspect page breaks again.
Access-denied or authentication error The URL is protected or the request lacks required credentials or headers. Use documented authentication or header controls only when authorized; do not assume a conversion option bypasses site access rules.
A local HTML conversion cannot read a local asset Local-file access is disabled or restricted. Use the manual’s narrow allow-list mechanism for required files; enable broader access only when necessary.
An option is rejected or behaves differently The installed package’s manual or build differs from the example. Run wkhtmltopdf --version and consult that installation’s help/manual.

10. Performance, reliability, and cost

wkhtmltopdf runs locally as a command-line conversion tool, so conversion time depends on the page, its remote resources, the selected waits, and the machine. A longer JavaScript delay can help a late-loading page but also extends the conversion. Avoid adding waits or options that the page does not need.

For repeatable archival, retain the source URL, capture date, command options, and resulting PDF together in your own records. Revisit important outputs after changing the installed build or layout settings. The tool’s documented controls do not establish universal rendering compatibility for modern dynamic pages, and a successful process exit alone is not proof that all content was captured.

The project describes wkhtmltopdf as open source and provides binary downloads or source builds; the dossier does not specify paid service costs, current release support, or maintenance status. Check the project’s documentation and your distribution’s terms and package details before choosing an installation.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. For a PDF capture, add format=pdf to the one-call request. See the ScreenshotNeo API documentation for the supported parameters.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and responses identify the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does wkhtmltopdf save a webpage as an image too?

This guide uses its PDF output. ScreenshotNeo supports PNG, JPEG, WebP, and PDF responses from its screenshot API.

Can I archive a page I do not have permission to access?

Authentication and custom headers are for pages you are authorized to access. They do not grant permission to protected content.

Should I use a delay on every conversion?

No. Add a delay only when the page needs time to render additional content; extra waiting increases conversion time and still may not make every dynamic page render correctly.

Sources