How to Convert a Web Page to PDF in JavaScript
Convert web pages to PDFs in JavaScript with browser print, Puppeteer, Playwright, CSS print rules, troubleshooting, and a hosted API option.
To convert a web page to PDF in JavaScript, render it in a real browser and call its PDF API. For Node.js automation, Puppeteer and Playwright both follow the same sequence: launch a browser, navigate to the URL, wait for the page and fonts, call page.pdf(), then close the browser. For a person already viewing the page, use the browser’s print workflow and let them choose Save as PDF.
Choose the right conversion method
| Need | Recommended approach | Reason |
|---|---|---|
| A person saves the current page | Browser print workflow | No server-side browser is needed, and the user’s session, fonts and settings are available. |
| A backend converts known URLs or templates | Puppeteer or Playwright | A programmable browser runs the page’s HTML, CSS and JavaScript before creating the PDF. |
| PDF bytes must be streamed onward | Playwright buffer or Puppeteer createPDFStream() |
The result can be written to storage or sent in an HTTP response without an intermediate file. |
| The PDF must look like the screen | Emulate screen media before export | Both libraries use print media by default. |
1. Let a user export the page from the browser
This is the lightest solution when a human is present. A button can open the browser print dialog:
<button type="button" id="print-page">Save as PDF</button>
<script>
document.querySelector('#print-page').addEventListener('click', () => {
window.print();
});
</script>
The user chooses the destination, paper size, margins, headers and footer in their browser. Exact output varies by browser and by the user’s settings, so this is not ideal for unattended jobs or repeatable documents.
Print CSS for user and automated exports
@media print {
nav,
.cookie-banner,
.chat-widget,
.print-button,
.interactive-controls {
display: none !important;
}
@page {
size: A4;
margin: 18mm;
}
a {
color: inherit;
text-decoration: none;
}
.avoid-break {
break-inside: avoid;
}
}
Keep print rules focused on document structure: remove navigation and controls, choose page margins, and prevent important cards or table rows from splitting where practical. If you need the screen stylesheet instead, select screen media in Puppeteer or Playwright.
2. Convert a page with Puppeteer
Install Puppeteer in a Node.js project:
npm install puppeteer
Create export-puppeteer.mjs:
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.goto('https://example.com', {
waitUntil: 'networkidle2',
timeout: 60_000
});
// Use this only when the PDF should match screen CSS rather than print CSS.
// await page.emulateMediaType('screen');
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
landscape: false,
margin: {
top: '18mm',
right: '18mm',
bottom: '18mm',
left: '18mm'
}
});
} finally {
await browser.close();
}
Puppeteer’s documented flow is launch, navigate, call page.pdf(), and close. Page.pdf() generates a PDF with the print CSS media type by default, and the PDF operation waits for fonts by default. See the Puppeteer PDF generation guide and the Page.pdf API.
Useful Puppeteer PDF options
| Option | Use |
|---|---|
path |
Write the PDF to a file. Omit it when you need a buffer or stream. |
format |
Use a standard paper size such as A4 or Letter. |
width, height |
Set a custom page size when a standard format is unsuitable. |
landscape |
Use horizontal orientation for wide tables or dashboards. |
margin |
Set top, right, bottom and left margins. |
printBackground |
Include background colors and images; set it explicitly for consistent output. |
pageRanges |
Export selected pages when a complete document is unnecessary. |
preferCSSPageSize |
Let CSS @page dimensions take precedence when your document defines them. |
displayHeaderFooter, templates |
Add controlled headers and footers when your document requires them. |
Return a PDF stream from Puppeteer
const stream = await page.createPDFStream({
format: 'A4',
printBackground: true
});
for await (const chunk of stream) {
process.stdout.write(chunk);
}
Puppeteer documents createPDFStream() for streaming PDF bytes. In an HTTP server, set Content-Type: application/pdf and pipe the stream to the response.
3. Convert a page with Playwright
Install Playwright and its browser binaries:
npm install playwright
npx playwright install chromium
Create export-playwright.mjs:
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const browser = await chromium.launch();
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 900 },
deviceScaleFactor: 1
});
await page.goto('https://example.com', {
waitUntil: 'networkidle',
timeout: 60_000
});
// Uncomment when the PDF should use screen media styles.
// await page.emulateMedia({ media: 'screen' });
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
landscape: false,
margin: {
top: '18mm',
right: '18mm',
bottom: '18mm',
left: '18mm'
}
});
await writeFile('page.pdf', pdf);
} finally {
await browser.close();
}
Playwright’s page API states that page.pdf() uses print CSS media by default. Call page.emulateMedia({ media: 'screen' }) to select screen media.
4. Wait for the page to be complete
Navigation completion does not always mean application data is ready. Single-page applications, lazy images, charts and API calls may finish after the initial document load.
await page.goto(url, { waitUntil: 'networkidle2', timeout: 60_000 });
await page.waitForSelector('[data-report-ready]', { timeout: 30_000 });
await page.evaluate(() => document.fonts.ready);
await page.waitForTimeout(300);
- Use a navigation wait condition appropriate to the site. Puppeteer commonly uses
networkidle2; Playwright supportsnetworkidle. - Wait for an application-specific selector such as
[data-report-ready]when data rendering has a known end state. - Wait for
document.fonts.readywhen font metrics affect wrapping or pagination. - For lazy content, scroll deliberately before export or make the application render print content without scrolling.
- Use a short, explicit delay only for a known animation or chart transition; a large arbitrary delay increases latency without guaranteeing readiness.
5. Authentication, images and browser context
Authenticated pages
Log in through the same browser context, set cookies before navigation, or attach an authorization header when the application supports it. Never put credentials in a public URL. Verify that the PDF job uses the intended tenant and user context.
Cross-origin images
Images must be reachable from the browser process. Check redirects, signed URL expiry and server responses. A page can look correct in a local browser while a server-side job receives an expired or blocked image URL.
Charts and client-side data
Wait for the chart’s rendered element or a data-ready marker. If the chart is animated, disable animation for print or wait until its final state. Capture a representative route in CI so a library or stylesheet change does not silently alter pagination.
6. cURL, Python and Node.js alternatives
The browser libraries above run inside your application. If you prefer a hosted conversion endpoint, ScreenshotNeo renders a URL and can return a PDF. Its API supports PDF paper size, margins, landscape mode and page ranges; the complete option names and examples are in the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF is blank | The export ran before client-side rendering finished, or navigation failed. | Check the response and console, increase the navigation timeout, and wait for a page-specific ready selector. |
| Fonts are wrong or text wraps differently | Fonts were not reachable or had not finished loading. | Verify font responses, await document.fonts.ready, and ensure the font URLs work from the server. |
| Background colors are missing | Background printing is disabled. | Set printBackground: true and check @media print rules. |
| Screen layout changes in the PDF | Print media is the default. | Call page.emulateMediaType('screen') in Puppeteer or page.emulateMedia({ media: 'screen' }) in Playwright. |
| Images are missing | Cross-origin, expired, blocked or lazy-loaded image URLs. | Inspect image responses, use valid URLs, and wait for the image or its containing component. |
| Content is cut off | Fixed heights, overflow rules or an unsuitable paper size. | Remove print-time fixed heights, review overflow, set a suitable format or use landscape mode. |
| Pages split in awkward places | Elements do not have print break rules. | Use break-inside: avoid, break-before and break-after where appropriate. |
| Navigation times out | Long polling, analytics or a page that never becomes idle. | Use a reliable application-ready selector instead of waiting indefinitely for network idle, and set a bounded timeout. |
| Browser cannot launch in production | Missing browser binary or sandbox/container dependencies. | Install the required Playwright browser or Puppeteer dependencies and follow your deployment platform’s process-user requirements. |
8. Performance, reliability and cost
- Reuse browsers carefully: launching Chromium for every request adds startup time. A controlled browser pool can reduce latency, while a fresh context per job keeps cookies and storage isolated.
- Bound every wait: set navigation, selector and overall job timeouts. Always close pages and browsers in a
finallyblock. - Control concurrency: too many simultaneous pages consume CPU and memory and can cause timeouts. Queue large batches and limit active browser contexts.
- Make output deterministic: pin browser versions, specify paper dimensions, margins, orientation and background printing, and disable animations in print CSS.
- Cache when inputs are stable: cache by URL plus relevant authentication, locale and document-version inputs. Invalidate when content changes.
- Measure the expensive steps: record navigation time, readiness wait, PDF generation time, output size and failure reason. This identifies whether the page or browser is the bottleneck.
- Hosted API economics: self-hosting trades per-request charges for browser infrastructure and maintenance. ScreenshotNeo bills only clean shots; bot checks, blank pages, timeouts, failed loads and cache hits cost nothing.
9. Security checklist
- Allow only intended URL schemes and hosts if users can submit URLs; otherwise a PDF worker can become a server-side request forgery path.
- Keep API keys, cookies and authorization headers out of logs and generated filenames.
- Use isolated browser contexts for different users and tenants.
- Set file-size and page-count limits for untrusted documents.
- Sanitize user-provided HTML and avoid executing untrusted JavaScript in a privileged context.
FAQ
Does JavaScript in the page run before the PDF is created?
Yes. Puppeteer and Playwright render the page in a browser, so client-side JavaScript can run. You still need an explicit readiness condition for asynchronous application data.
Why does the PDF use print styles?
Both libraries select print CSS media by default. Emulate screen media when the screen layout is the required output.
Can I convert a page without Node.js?
Yes. A browser print dialog works for a human user, and a hosted API such as ScreenshotNeo can accept a URL from cURL or Python.
Should I use Puppeteer or Playwright?
Either is a sound starting point. Choose based on the browser support, deployment tooling and test stack already used by your project; both expose the core page.pdf() workflow.
How do I create a PDF from HTML that is not publicly reachable?
Render it inside your authenticated application or serve it to the browser worker on a private, controlled route. Pass the required session context securely and restrict who can request the job.


