ScreenshotNeo

BlogHTML to image & PDF

How to Convert an HTML Link to PDF

Convert any HTML link to PDF manually or automatically with Chrome, Playwright, wkhtmltopdf, and ScreenshotNeo.

By the ScreenshotNeo team1 October 20267 min read

The quickest way to convert an HTML link to PDF is to open the link in a browser, choose Print, select Save as PDF, and save the file. For repeatable or server-side conversion, use headless Chrome, Playwright, wkhtmltopdf, or a hosted API such as ScreenshotNeo.

Choose the right HTML-to-PDF method

Use case Best method Why
One page, occasional use Browser Print No installation or code required
Shell script or CI job Headless Chrome Simple URL-to-PDF command
Application automation Playwright Control over waiting, authentication, CSS, pages, and output
Legacy command-line workflow wkhtmltopdf Accepts URLs and local HTML files
Managed conversion without browser setup ScreenshotNeo One HTTP request, PDF options, clean captures, and usage-based billing
  1. Open the HTML link in Chrome, Edge, Firefox, or another browser.
  2. Wait until the content you need is visible. Expand sections or sign in if required.
  3. Open the browser menu and select Print, or press Ctrl+P on Windows/Linux or ⌘+P on macOS.
  4. Choose Save as PDF as the destination.
  5. Set paper size, orientation, margins, scale, background graphics, and page range.
  6. Save the PDF.

This route works for a webpage URL and for a local .html file after opening it in a browser. The PDF can differ from the screen because browsers apply print-specific CSS.

Browser print settings to check

  • Background graphics: enable this when colors, panels, or images are part of the document.
  • Margins: use default margins for ordinary documents; choose none or custom margins for edge-to-edge layouts.
  • Scale: reduce scale when content is clipped or spills onto unexpected pages.
  • Headers and footers: disable them when you do not want the URL, date, or page number printed.
  • Selection and page range: print only the required pages when the document is long.

Chrome supports a command-line PDF export with --headless and --print-to-pdf. The default output filename is usually output.pdf. Add --no-pdf-header-footer when you want to omit generated headers and footers. See the Chrome headless documentation.

google-chrome --headless --disable-gpu \
  --no-pdf-header-footer \
  --print-to-pdf=page.pdf \
  "https://example.com/article"

On systems where the executable is named differently, use chromium, chromium-browser, or the installed Chrome path.

Convert a local HTML file with Chrome

google-chrome --headless --disable-gpu \
  --no-pdf-header-footer \
  --print-to-pdf=local.pdf \
  "file:///absolute/path/to/page.html"

Use an absolute file:// URL. Relative assets may fail if the document expects a web server; serving the directory over HTTP is often more reliable.

Headless Chrome limitations

  • The command does not provide application-level control over login flows, custom waits, or interactions.
  • Pages that load data after the initial navigation may be captured before the data appears.
  • Print CSS controls the output, so a screen layout may not match the PDF.

Playwright automates a real browser and exposes page.pdf(). PDF generation is supported in Chromium. By default, PDF output uses print CSS media. If the page needs screen styles, call page.emulateMedia({ media: 'screen' }) before generating the PDF. Refer to the Playwright page.pdf() API.

Install Playwright

npm install playwright
npx playwright install chromium

Complete Node.js example

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage({
    viewport: { width: 1440, height: 900 },
  });

  await page.goto('https://example.com/article', {
    waitUntil: 'networkidle',
    timeout: 90_000,
  });

  // Use this when the screen stylesheet should control the PDF.
  // await page.emulateMedia({ media: 'screen' });

  await page.pdf({
    path: 'article.pdf',
    format: 'A4',
    printBackground: true,
    margin: {
      top: '16mm',
      right: '16mm',
      bottom: '16mm',
      left: '16mm',
    },
    displayHeaderFooter: false,
    preferCSSPageSize: true,
  });

  await browser.close();
})();

Useful Playwright PDF options

Option Purpose
path Writes the PDF to a file.
format Sets a standard paper size such as A4 or Letter.
width, height Sets custom page dimensions.
landscape Uses landscape orientation.
printBackground Includes background colors and images.
margin Sets top, right, bottom, and left margins.
pageRanges Exports selected pages, such as 1-3.
preferCSSPageSize Lets CSS @page size take priority.
displayHeaderFooter Controls generated header and footer templates.

Wait for dynamic content before exporting

await page.goto('https://example.com/dashboard', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.report-ready', { timeout: 30_000 });
await page.waitForTimeout(500);
await page.pdf({ path: 'dashboard.pdf', format: 'A4', printBackground: true });

Prefer a meaningful selector such as .report-ready over an arbitrary long delay. For authenticated pages, create a context with the required cookies or complete the login flow before calling page.pdf().

Method 4: Convert a URL or file with wkhtmltopdf

wkhtmltopdf accepts a webpage URL or an HTML filename and writes a PDF. Its documented rendering engine is Qt WebKit, so compare its output with a Chromium-based method when modern CSS or JavaScript matters. See the wkhtmltopdf usage documentation.

wkhtmltopdf "https://example.com/article" article.pdf
wkhtmltopdf "/absolute/path/to/page.html" local.pdf

Because rendering engines differ, pages designed for current Chromium may have layout, font, or JavaScript differences in wkhtmltopdf.

Or skip the browser setup

ScreenshotNeo converts a URL to a PDF through one GET request. Its capture options include paper size, margins, landscape mode, and page ranges. Before capture, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result in X-Page-Verdict and X-Billed headers. An MCP server also lets Claude, Cursor, and other MCP clients call screenshot and PDF tools.

See the ScreenshotNeo API documentation for all parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = require('node:fs');
fs.writeFileSync('shot.pdf', Buffer.from(await res.arrayBuffer()));

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

PDF layout controls that affect the result

Browsers can apply a print stylesheet, including @media print rules and @page declarations. A navigation bar, background color, or interactive widget may therefore appear differently or disappear in the PDF. In Playwright, use page.emulateMedia({ media: 'screen' }) when the screen stylesheet is the intended design.

Backgrounds and images

Enable background printing in the browser or set printBackground: true in Playwright. Wait for important images and fonts before export. Lazy-loaded images may require scrolling or an application-specific readiness selector.

Page size, margins, and page breaks

Use a consistent paper size and explicit margins for documents that will be printed. CSS rules such as break-inside: avoid, break-before, and @page can reduce awkward splits, but the final result depends on the rendering engine.

Troubleshooting HTML-to-PDF conversion

Symptom Likely cause Fix
Blank PDF Navigation failed, content requires JavaScript, or a bot check appeared Check the URL in a normal browser, wait for a readiness selector, and inspect response status and page content.
Missing charts or images Resources were still loading or backgrounds were disabled Wait for the relevant selector or network idle, then enable background printing.
Wrong layout Print CSS differs from screen CSS Inspect @media print; in Playwright emulate screen media when appropriate.
Content is cut off Viewport, paper size, scale, or fixed-width CSS is too large Set the intended paper size, adjust margins or scale, and inspect responsive breakpoints.
Only the first page exports The page is still loading or the tool is capturing a viewport instead of a PDF Use a PDF-specific command or page.pdf(), and wait for the document to finish rendering.
Fonts differ Web fonts were unavailable at capture time Wait for document.fonts.ready in automation and verify the font URL is reachable.
Login page appears The converter has no authenticated session Supply cookies or authorization in an automated browser context, or convert after signing in manually.
wkhtmltopdf renders modern CSS incorrectly Its Qt WebKit engine differs from current Chromium Use headless Chrome or Playwright when browser fidelity is more important.
Chrome command is not found Chrome is not installed or has another executable name Install Chromium/Chrome or call the full executable path.

Performance, reliability, and cost notes

  • Manual printing: lowest setup cost for a single page, but difficult to repeat consistently.
  • Headless Chrome: suitable for scripts and CI; keep the browser version fixed when reproducibility matters.
  • Playwright: reuse a browser process for batches, set navigation and selector timeouts, and close contexts after each job.
  • wkhtmltopdf: simple to run, but validate output against the page’s CSS and JavaScript requirements.
  • Dynamic pages: wait for a deterministic application signal rather than relying only on a fixed delay.
  • Hosted conversion: removes browser installation and maintenance. ScreenshotNeo offers caching with a chosen TTL, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, and a usage API. Failed loads and other non-clean results are not billed.

Conversion checklist

  • Confirm the URL is publicly reachable or provide an authenticated session.
  • Choose print or screen media intentionally.
  • Set paper size, orientation, margins, and page range.
  • Enable backgrounds when visual styling matters.
  • Wait for dynamic data, images, and fonts.
  • Check page breaks and the final PDF at its intended size.
  • For repeated jobs, record failures and use timeouts and retries.

FAQ

Can I convert a local HTML file to PDF?

Yes. Open it in a browser and print it, pass its absolute file:// URL to headless Chrome, or provide its filename to wkhtmltopdf.

Which method preserves the webpage exactly?

No method guarantees an identical screen and PDF because print CSS, page dimensions, fonts, and pagination affect output. Chromium-based Playwright or Chrome is usually the most direct match for pages built for modern browsers.

Can a PDF include content loaded after page load?

Yes, if the converter waits until that content is rendered. Use a readiness selector or application signal in automation.

Playwright and command-line scripts can process multiple URLs. ScreenshotNeo also supports bulk capture for up to 100 URLs per call.

Does ScreenshotNeo return a PDF or an image?

Its screenshot API can return PNG, JPEG, WebP, or PDF; select the PDF output and configure paper and layout options as needed.