ScreenshotNeo

BlogHTML to image & PDF

How to Convert a Web Page to PDF in Bun

Use Puppeteer or Playwright from Bun to print a web page to PDF, with working code, print options, readiness checks, and fixes for common failures.

By the ScreenshotNeo team29 September 202611 min read

How to Convert a Web Page to PDF in Bun

To convert a web page to PDF in Bun, use Bun to run Puppeteer or Playwright, navigate to the page, wait for its content to be ready, and call page.pdf(). Bun is the JavaScript runtime in this setup; Puppeteer or Playwright supplies the browser automation and PDF API. The result is a PDF file you can save or return from an application. Puppeteer’s documented flow uses page.goto(), page.pdf(), and browser.close(), and waits for web fonts by default. Puppeteer PDF generation.

This guide uses Puppeteer for the main example because its PDF workflow is direct and documented. Playwright is also a good choice when it is already part of your project. Bun.WebView can drive a browser, but Bun documents it as experimental and does not provide a high-level page.pdf() equivalent; its raw Chrome DevTools Protocol route is more version-sensitive.

1. Install Puppeteer and run the first PDF script

Create a project, install Puppeteer, then save the following as pdf.ts. Puppeteer’s installation process can download a compatible Chrome browser. If your package manager has blocked install scripts, install the browser explicitly as described below. See the official installation guide.

bun init -y
bun add puppeteer

The script accepts a URL as its first argument and writes page.pdf to the current directory. It also closes the browser if navigation or PDF generation fails.

import puppeteer from 'puppeteer';

const url = process.argv[2] ?? 'https://example.com';
const browser = await puppeteer.launch();

try {
  const page = await browser.newPage();
  await page.goto(url, {
    waitUntil: 'networkidle2',
    timeout: 45_000,
  });

  await page.pdf({
    path: 'page.pdf',
    format: 'A4',
    printBackground: true,
  });

  console.log('Saved page.pdf');
} finally {
  await browser.close();
}

Run it with Bun and provide a page URL:

bun run pdf.ts https://example.com

page.pdf() returns PDF bytes if you omit path; write them with Bun’s file API when you need to control the output location or return the bytes from a server route:

const pdf = await page.pdf({ format: 'A4', printBackground: true });
await Bun.write('page.pdf', pdf);

2. Wait for the right page state

Navigation completion and application readiness are different things. networkidle2 waits until network activity is low, which often works for ordinary pages. A page with analytics, polling, streaming, or a persistent connection may never become idle. In those cases wait for the specific content your PDF needs instead of treating network idleness as proof that the page is finished.

A reliable conversion waits for the page’s content before the browser prints it.
A reliable conversion waits for the page’s content before the browser prints it.

For example, if a single-page app renders a report element after loading its data, use a selector wait. Replace the selector with one that appears only when the desired content is present:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 20_000 });
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

Other useful strategies:

  • domcontentloaded is a faster navigation milestone; follow it with an explicit selector or application readiness condition.
  • load waits for the page load event and its dependent resources, but client-side rendering can continue after that event.
  • networkidle2 is a practical default for many pages, but may not suit sites with ongoing requests.
  • A fixed delay can help when a page has a known animation or delayed widget, but it is less reliable than waiting for a meaningful selector or state.

Puppeteer waits for web fonts during PDF generation by default. That does not guarantee that application data, lazy-loaded images, charts, or animations have reached the state you want. For lazy content, scroll through the page before printing or trigger the site’s own load-more behavior, then check the relevant elements.

3. Choose page size, margins, and print appearance

Puppeteer’s page.pdf() prints using the CSS print media type by default. A site can therefore hide navigation, change layout, or use print-specific styles. To generate the screen appearance instead, call page.emulateMediaType('screen') before page.pdf(). Playwright has the analogous page.emulateMedia({ media: 'screen' }) method. See the Puppeteer Page API and Playwright Page API.

await page.emulateMediaType('screen');
await page.pdf({
  path: 'screen-style.pdf',
  format: 'A4',
  printBackground: true,
  margin: { top: '12mm', right: '12mm', bottom: '16mm', left: '12mm' },
});

Common PDF options include:

Option What it controls When to use it
format Named paper size, such as A4 or Letter Use for ordinary documents with a known paper standard.
width, height Explicit page dimensions Use for custom formats; provide dimensions with units such as mm or px.
margin Top, right, bottom, and left whitespace Set explicit margins when the browser defaults or page CSS do not fit.
printBackground Whether background colors and images are printed Set true for branded pages, colored panels, and background artwork.
preferCSSPageSize Whether CSS @page size takes precedence Use when the site’s print stylesheet defines the intended sheet size.
displayHeaderFooter Browser header and footer templates Enable when printed pages need page numbers, a title, or a date.
landscape Landscape orientation Use for wide tables, dashboards, or diagrams.
pageRanges Pages to include Use to extract a subset, such as 1-3.

For exact colors, add print CSS such as -webkit-print-color-adjust: exact to the relevant elements. Print CSS is usually the best place to hide interactive controls, set page breaks, and prevent awkward splits. For example:

@media print {
  nav, .cookie-banner, .screen-only { display: none !important; }
  article { break-inside: avoid; }
  h1, h2 { break-after: avoid; }
  body { -webkit-print-color-adjust: exact; }
}

4. A more production-ready Puppeteer function

In an application, accept a URL only from a trusted source or validate it before navigation. A browser can request resources from the network, so blindly passing user-controlled URLs into a renderer can expose internal services. Set finite navigation and readiness timeouts, use a unique output path when writing files concurrently, and always close the browser. For higher throughput, consider reusing a browser process and opening a fresh page per job, while ensuring pages and contexts are closed after use.

import puppeteer from 'puppeteer';

export async function renderPdf(url: string): Promise {
  const parsed = new URL(url);
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error('Only http and https URLs are supported');
  }

  const browser = await puppeteer.launch();
  try {
    const page = await browser.newPage();
    page.setDefaultNavigationTimeout(45_000);
    await page.goto(parsed.href, { waitUntil: 'networkidle2' });
    // If this page renders asynchronously, wait for its own ready selector here.
    return await page.pdf({
      format: 'A4',
      printBackground: true,
      margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
    });
  } finally {
    await browser.close();
  }
}

const bytes = await renderPdf(process.argv[2] ?? 'https://example.com');
await Bun.write('page.pdf', bytes);

This example launches a new browser for clarity. Browser startup and memory use can be significant compared with the PDF call itself. For a service that handles repeated jobs, measure the workload and consider a managed browser lifecycle, concurrency limits, and job timeouts. Do not let unbounded requests create unbounded browser processes.

5. Use Playwright instead

Playwright also runs from Bun and exposes page.pdf() in its Chromium browser. Its PDF output uses print CSS by default; emulate screen media when needed. Install the library and browser binaries according to its official setup guide, since package and browser installation steps can vary by environment.

import { chromium } from 'playwright';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle', timeout: 45_000 });
  await page.pdf({
    path: 'page.pdf',
    format: 'A4',
    printBackground: true,
    margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
  });
} finally {
  await browser.close();
}

Choose based on the surrounding project and the browser features you need. Both libraries give you browser control, page readiness waits, and PDF settings. Puppeteer’s guide offers a particularly compact PDF recipe; Playwright may fit naturally into an existing Playwright automation setup. In either case, the browser executable and system dependencies are part of deployment, not supplied by Bun itself.

6. What about Bun.WebView?

Bun.WebView is Bun’s experimental browser API. It can navigate, execute JavaScript, interact with the page, and take screenshots. On macOS it can use system WebKit; its Chrome backend can drive an installed Chrome-family browser over the Chrome DevTools Protocol on macOS, Linux, and Windows. The Chrome backend searches for browser executables and can also use Playwright’s chrome-headless-shell cache. Check the current Bun reference for exact requirements before relying on it.

The documented WebView reference exposes a raw cdp(method, params?) bridge after navigation, but does not document a high-level PDF method. Chrome DevTools Protocol supports print-to-PDF operations, so calling that through the bridge may be possible in a specific Bun and backend version. Treat it as an advanced integration: it is less writer-friendly than page.pdf(), depends on Chrome mode, and should be pinned and verified against the Bun release you deploy. For a stable tutorial or general PDF service, use Puppeteer or Playwright.

7. Troubleshooting common failures

Symptom Likely cause Fix
Could not find Chrome Puppeteer’s browser download was skipped or the cache is missing. Run bun x puppeteer browsers install, or configure an explicit executable if you manage Chrome yourself. See Puppeteer installation.
Navigation times out The site is slow, has persistent requests, or never reaches the chosen idle condition. Use domcontentloaded and wait for a specific selector, increase a finite timeout, and inspect whether the URL is reachable from the host.
PDF is blank or missing content The app has not rendered data, a redirect/auth flow occurred, or the wrong frame/state was printed. Wait for a page-specific ready selector; check the final URL and status; provide required authentication where appropriate.
Colors or backgrounds are absent Print styling alters colors or PDF defaults omit backgrounds. Set printBackground: true and use -webkit-print-color-adjust: exact where appropriate.
Layout differs from the browser view Print CSS is active by default or the paper dimensions trigger responsive breakpoints. Inspect @media print rules; use screen media only if the screen layout is the intended output; set paper dimensions and margins explicitly.
Images or fonts are missing Resources are delayed, blocked, lazy-loaded, or inaccessible from the runtime. Wait for the needed assets or application state; scroll to trigger lazy loading; verify network access and avoid printing before content is ready.
Browser fails to launch in a container The image lacks a compatible browser binary or system libraries, or sandbox settings differ. Use a deployment image with the browser and its dependencies installed; follow the library’s official container guidance and the host’s security requirements.
WebView cdp() call fails The view has not navigated, the Chrome backend is unavailable, or the command differs across versions. Navigate first, confirm Chrome mode and executable availability, and check the CDP method against the target Bun release.

8. Performance, reliability, and cost

PDF generation is browser work: navigation, scripts, fonts, images, layout, and PDF encoding all take time and memory. Reduce avoidable work by waiting for a meaningful readiness signal instead of a long fixed delay, blocking irrelevant heavy resources only when you know they are not needed, and limiting concurrent pages. Reuse browser processes in a long-running worker if startup cost matters, but close each page or context and recycle browsers according to the needs of your deployment.

Removing overlays before capture can keep them out of the resulting document.
Removing overlays before capture can keep them out of the resulting document.

For reliable output, keep browser and automation-library versions compatible, install the browser during image or environment setup, and handle timeouts and navigation errors as failed jobs. Use a job queue for bursts rather than launching unlimited concurrent conversions. Record the input URL, final URL, elapsed time, and failure category without logging credentials or sensitive page contents. Validate the output bytes as a PDF before publishing or attaching it.

There is no single cost for “Bun PDF conversion”: costs depend on where the browser runs, its CPU and memory, concurrency, and the pages rendered. Self-hosting gives control but makes browser installation, patching, capacity, and operations your responsibility. A hosted capture API moves browser setup out of your application and charges according to that service’s plan and billing rules.

Or skip the browser setup

If you need a PDF of a URL without installing and operating a browser, ScreenshotNeo is a website screenshot and PDF API. It can also return PNG, JPEG, or WebP. A single GET request accepts the URL and PDF options; see the ScreenshotNeo API documentation for the current parameters.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -d format=pdf \
  -o page.pdf

From Python, the same URL-based request can be made with requests:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)

Or from Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
  format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await Bun.write('page.pdf', new Uint8Array(await res.arrayBuffer()));

ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. It also provides an MCP server for AI agents, including Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

FAQ

Can Bun convert HTML strings to PDF?

Yes, with a browser library: open a page, call page.setContent(html), and then call page.pdf(). External fonts, images, and stylesheets still need to resolve, so wait for the content and assets the document depends on.

Does Bun include a browser for PDF generation?

Bun itself is the runtime. Puppeteer or Playwright controls a browser installation; Bun.WebView is an experimental alternative with platform-specific browser requirements.

Why does the PDF use different CSS than the website?

PDF generation defaults to print media, where sites commonly define different layout rules. Check the page’s print stylesheet or emulate screen media if that output is specifically required.

Can I generate PDFs from authenticated pages?

Yes, if the browser session has valid access. Configure cookies or authentication through the chosen automation library and keep credentials out of source control and logs.

Is WebView’s CDP printing a drop-in replacement for page.pdf()?

No. Bun documents a raw CDP bridge rather than a high-level PDF method. Its use depends on backend and Bun versions, so verify the exact integration before adopting it.