ScreenshotNeo

BlogHow-to

Convert HTML to PDF with Clickable Links Using Playwright

Use Playwright’s Chromium-based page.pdf() to save HTML as PDF, control print layout, and check whether links work in your target PDF viewer.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright’s page.pdf() method with Chromium to turn a rendered HTML page into a PDF. Keep ordinary anchor elements and their href values in the source, then open the resulting PDF in the viewer your users rely on and check that both external and in-document links work. The Playwright API documents PDF generation with print CSS media by default; it does not explicitly guarantee that every link becomes a working PDF annotation. Playwright Page API

1. Install Playwright and Chromium

PDF export through Playwright is Chromium-only. Install Playwright and its Chromium browser binary in the environment that will generate the PDF. If you deploy on a Linux host or container, install the browser’s OS dependencies too. Playwright recommends keeping the package and browser installations updated together. Playwright PDF export · Playwright browsers

npm init -y
npm install playwright
npx playwright install chromium
# On Linux, install the Chromium OS dependencies as needed:
npx playwright install-deps chromium

Save the following as html-to-pdf.js. It navigates to a page, waits for navigation to reach the chosen state, writes page.pdf to disk, and closes Chromium even if an operation fails.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', {
      waitUntil: 'networkidle',
      timeout: 30_000,
    });

    await page.pdf({
      path: 'page.pdf',
      format: 'A4',
      printBackground: true,
    });
    console.log('Wrote page.pdf');
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with node html-to-pdf.js. This is an illustrative implementation. Choose a readiness condition that suits the page: networkidle can be a poor fit for applications with ongoing network activity, while load may happen before client-side content or lazy images are ready. For a known app, wait for a meaningful selector or application-specific ready signal before exporting.

Use semantic anchors with real href attributes. An anchor provides the source hyperlink information; it does not by itself prove that a PDF renderer emitted a clickable PDF annotation. WHATWG HTML Standard

<a href="https://example.org/docs">Read the documentation</a>

<a href="#appendix">Jump to the appendix</a>

<h2 id="appendix">Appendix</h2>

For content you supply directly, set it with page.setContent() and wait for fonts or other required assets before calling page.pdf(). For a remote page, confirm the browser reached the intended final URL and state. Links created only by a later interaction or script need that interaction or script to run before PDF generation.

const html = `<!doctype html>
<html>
  <head><meta charset="utf-8"></head>
  <body>
    <p>See <a href="https://example.org">the reference</a>.</p>
    <p><a href="#end">Go to the end</a></p>
    <div style="height: 1200px"></div>
    <h2 id="end">End</h2>
  </body>
</html>`;

await page.setContent(html, { waitUntil: 'load' });
await page.pdf({ path: 'supplied.html.pdf', format: 'A4' });

3. Choose print or screen styling

page.pdf() renders using print CSS media by default. That means rules inside @media print and @page can change the page’s appearance or pagination. If the intended output should follow screen styles, explicitly emulate screen media before export:

await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-styled.pdf', format: 'A4', printBackground: true });

Use screen emulation only when screen styling is the desired document design. If the output is intended for printing, leave print media active and adjust the print stylesheet. Inspect both @media print and @page when colors, visibility, columns, or page breaks differ from the browser view.

4. Adjust paper size, margins, and pagination

The documented defaults and main layout options matter when a PDF looks clipped or paginates unexpectedly. Check the current Page API reference when upgrading Playwright.

Option Purpose Practical guidance
format Named paper size; the documented default is Letter. Set A4, Letter, or another supported paper format explicitly for predictable output.
width, height Paper dimensions when defining a custom page size. Use a supported CSS length such as mm, in, or px. Avoid relying on the default when a specific stock size is required.
margin Space around the printed page content. Set each side deliberately when a layout needs room for content or print-safe edges.
pageRanges Limits which pages are included. Useful for extracting selected pages from a longer rendered document.
scale Scales the rendered content; default is 1. Use small adjustments only after checking paper size and CSS. Scaling down can make text too small.
printBackground Includes background graphics; default is false. Set true when backgrounds or colored blocks are part of the intended design.
preferCSSPageSize Whether CSS @page size takes priority; default is false. Set true when the stylesheet owns the paper size and you want Chromium to honor it.
outline, tagged Optional PDF outline and tagged output; documented defaults are false. Enable only when the resulting document needs these structures, and inspect the output in your downstream tools.
path Writes the PDF to a file. Omit it to receive the PDF as a buffer instead.
await page.pdf({
  path: 'report.pdf',
  format: 'A4',
  landscape: false,
  printBackground: true,
  margin: {
    top: '15mm',
    right: '12mm',
    bottom: '15mm',
    left: '12mm',
  },
  scale: 1,
  preferCSSPageSize: true,
  pageRanges: '1-3',
});

When CSS defines the page size, for example with @page { size: A4; margin: 15mm; }, set preferCSSPageSize: true if that CSS should control the PDF’s paper dimensions. Otherwise choose the size in Playwright and use CSS for the page content. Avoid competing size rules unless you have checked the resulting output.

5. Get a PDF buffer or serve it from an application

page.pdf() returns a PDF buffer. Use path to write directly to a file; without a path, you can send the bytes to a storage layer or HTTP response. Keep browser lifetime bounded and close it after each job or batch.

const pdf = await page.pdf({ format: 'A4', printBackground: true });

// Example for an Express route after page setup:
res.type('application/pdf');
res.setHeader('Content-Disposition', 'attachment; filename="page.pdf"');
res.send(pdf);

For user-provided HTML or URLs in a service, validate inputs and constrain where the browser can navigate according to your application’s security requirements. PDF generation loads page resources and can consume substantial memory for large documents; impose job timeouts and size limits appropriate to your service.

Do not infer link behavior solely from the HTML source or from a successful PDF write. Open the file in the target PDF viewer and test:

  1. An external HTTPS link opens the intended destination.
  2. An in-document fragment link jumps to the intended section.
  3. Links remain aligned with their visible text after pagination and scaling.
  4. Links on pages with print-specific layout changes still point to the right destinations.

If a link is missing or inactive, first confirm the anchor existed in the final rendered DOM before export. Then compare another viewer and inspect the PDF’s link annotations with the PDF tooling in your own workflow. The Playwright API reference documents PDF generation and layout options but does not expressly guarantee clickable link annotations, so viewer-level validation belongs in your release checks.

7. Troubleshooting

Symptom Likely cause Fix
page.pdf is not a function or PDF export fails under a non-Chromium engine The documented Playwright PDF export path is Chromium-only. Launch chromium for PDF generation. Do not assume WebKit or Firefox provides this same export capability.
Browser executable missing The Playwright package is installed but its browser binary is not. Run npx playwright install chromium in the deployment environment and keep the browser version aligned with the installed Playwright package.
Chromium fails to launch on Linux Required OS libraries may be absent. Install the required dependencies, for example with npx playwright install-deps chromium, or use the dependency setup appropriate to the host image.
PDF appearance differs from the browser PDF export uses print media by default; print CSS, @page, paper size, or margins may alter layout. Inspect print rules. Emulate screen only if screen-styled output is intended. Set paper and margins explicitly.
Colors or background panels disappear printBackground defaults to false. Set printBackground: true and confirm the print stylesheet does not remove those elements.
Page content is missing or stale The page was printed before client rendering, images, or fonts completed. Wait for a meaningful application-ready selector or required assets; select a navigation wait condition based on the page’s behavior.
Navigation times out on an app that keeps making requests A network-idle condition may never occur on pages with persistent traffic. Use an appropriate earlier navigation condition, then wait for a specific ready state rather than waiting indefinitely for all network activity to stop.
Links appear but do not activate The viewer may not recognize the generated annotations, or the final DOM may lack a usable anchor. Check the rendered <a href>, inspect the file in the target viewer, and validate both external and fragment links after export.
Content is clipped or too small Paper dimensions, margins, scaling, or wide content conflict. Choose the intended paper size, adjust margins or layout CSS, and use scale cautiously. Check a representative multi-page output.

8. Performance, reliability, and cost considerations

PDF generation runs a browser and lays out the full printable document, so time and memory needs depend on the page, its assets, and document length. Reuse a launched browser for multiple jobs when appropriate, but create an isolated page or context per job and close resources reliably. Avoid an arbitrary fixed sleep when a specific rendered state can be awaited. For large pages, test representative input sizes and set timeouts, concurrency, and resource limits for your environment.

Playwright is an open-source browser automation library; the browser work runs in the environment you operate. Budget for the compute, storage, and maintenance of browser binaries and OS dependencies. Keep Playwright and its browser installation current as recommended in the browser documentation. No throughput or cost benchmark is implied here: measure with your own pages and deployment limits.

Or skip the browser setup

If you need a page screenshot or PDF without maintaining a browser capture environment, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns a screenshot or PDF; here is the documented request shape for an image capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Its MCP server provides screenshot tools for AI agents, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. For clickable links in a PDF, inspect the returned PDF in your target viewer just as you would with a browser-generated file.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Can Playwright export a PDF with Firefox or WebKit?

The documented Playwright PDF export capability is Chromium-only. Playwright supports other engines for browser automation, but that does not make PDF export uniform across them. PDF export documentation

Why does my PDF not look like the page on screen?

page.pdf() uses print CSS media by default. Review @media print and @page, or emulate screen media if that is the intended output.

No guarantee is stated in the consulted Playwright API reference. Preserve normal anchors and verify the exported PDF in the target reader.

Can I generate only selected pages?

Yes. Use the pageRanges PDF option to select pages, and check the current API reference for its accepted syntax.