Puppeteer vs Playwright for Saving Webpages as PDF Archives
Compare Puppeteer and Playwright for saving webpages as PDFs, with runnable code, print layout options, troubleshooting, and practical archiving limits.
Short answer: Puppeteer and Playwright can both save a webpage as a PDF. Both generate print-media output by default; switch to screen media explicitly if you want screen styles. Choose based on your project’s browser automation needs and API preferences. The documentation reviewed does not establish that either tool is faster or produces more faithful PDFs.
A PDF is a rendered representation of a page, not necessarily a complete web archive. These APIs do not, by themselves, establish that the original HTML, linked assets, response headers, metadata, or a replayable browser environment have been preserved. If archival completeness matters, define what must be retained and verify it separately.
1. Choose Puppeteer or Playwright
| Question | Puppeteer | Playwright |
|---|---|---|
| What media does PDF use? | Print by default. | Print by default. |
| How do I save a file? | Pass a path to page.pdf(); the PDF guide demonstrates this pattern. |
Pass path to save to disk. Without a path, the API returns PDF bytes without saving a file. |
| What layout controls are documented in the reviewed sources? | The guide demonstrates the basic PDF workflow. Consult the current API reference for the exact options you need. | The API documents paper format or dimensions, margins, CSS page sizing, background graphics, scaling, and page ranges. |
| Which browser engines are documented? | The reviewed PDF sources do not establish a broader engine comparison. | Playwright documents Chromium, WebKit, Firefox, and branded Chrome and Edge channels. |
| Which is faster or more faithful? | The reviewed documentation provides no comparative benchmark for speed or fidelity. Measure your own pages and environment if those properties decide the choice. | |
Playwright’s documented engine breadth can matter if your automation project needs engines beyond Chromium. It does not establish identical PDF output across engines. For a Chromium-oriented workflow, either library can implement the basic navigate-and-save task.
2. Save a page as PDF with Puppeteer
Install Puppeteer in a Node.js project, then save a page by navigating, generating a PDF, and closing the browser. This example uses a navigation lifecycle option and writes to page.pdf.
npm install puppeteer
// save-pdf-puppeteer.js
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle2' });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
} finally {
await browser.close();
}
})();
Run it with node save-pdf-puppeteer.js. Puppeteer’s PDF guide says PDF generation waits for fonts by default. The example uses printBackground so background graphics are included; print color handling is discussed below.
3. Save a page as PDF with Playwright
Install Playwright and its browser binaries for the engine you plan to use. This Chromium example writes the PDF to disk.
npm install playwright
npx playwright install chromium
// save-pdf-playwright.js
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' }
});
} finally {
await browser.close();
}
})();
Run with node save-pdf-playwright.js. The example opts into CSS page sizing and sets explicit margins. Remove those options if the document’s CSS should not control paper size or if default margins suit the output.
4. Control print and screen rendering
Both APIs use print CSS media for PDF generation by default. Print media can hide navigation, rearrange layouts, or change colors compared with what a visitor sees on screen.
Use screen media when the PDF should resemble the screen layout
For Puppeteer, emulate screen media before calling page.pdf(). For Playwright, the documented pattern is page.emulateMedia({ media: 'screen' }) before PDF generation.
// Puppeteer: before page.pdf(...)
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-layout.pdf', printBackground: true });
// Playwright: before page.pdf(...)
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-layout.pdf', printBackground: true });
Media emulation changes which CSS rules apply; it does not guarantee a pixel-identical screen capture. If you need a screenshot image rather than a PDF print rendering, use the library’s screenshot API or a screenshot service.
Preserve colors and backgrounds
Print output may modify colors. If exact CSS colors matter, add -webkit-print-color-adjust: exact to the page’s print styling. Set the PDF option that includes background graphics as well; otherwise background colors and images may be omitted.
@media print {
html {
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
}
5. Set page size, margins, and page ranges
Playwright’s documented PDF options include paper format or explicit dimensions, margins, whether CSS @page sizing takes precedence, background graphics, scale, and page ranges. For example:
await page.pdf({
path: 'selected-pages.pdf',
format: 'Letter',
printBackground: true,
scale: 0.95,
pageRanges: '1-3, 5',
margin: { top: '0.5in', right: '0.5in', bottom: '0.5in', left: '0.5in' }
});
Use format for a standard paper size, or explicit width and height when the document needs custom dimensions. Use preferCSSPageSize when the site’s @page rules should determine the page size. Check the current Puppeteer API reference for the exact option names and support in the version you install; the reviewed Puppeteer guide focuses on the basic workflow and does not establish a complete option-by-option comparison.
6. Make the output more reliable
- Wait for a meaningful page state. Navigation completion does not always mean client-rendered content or images are ready. If a known element signals readiness, wait for it before generating the PDF.
- Choose network-idle settings carefully. Pages with analytics, polling, or long-lived requests may never become idle. Prefer a known selector or a bounded wait when network activity does not settle.
- Keep browser cleanup in a
finallyblock. This helps release browser processes when navigation or PDF generation fails. - Test representative pages. Long pages, pages with web fonts, dynamic content, print-specific styles, and large background images can produce different results.
- Record the rendering environment for repeatability. Browser version, installed fonts, page state, and network-loaded assets can affect rendered output. The PDF call alone does not preserve this environment.
7. Understand what “PDF archive” means
A PDF is useful for reading, sharing, and retaining a fixed rendering. It is not automatically a complete archival package. Before treating it as an archive, decide whether you also need the original HTML, linked files, source URL, capture time, response headers, metadata, or enough environment details to reproduce the rendering. The reviewed Puppeteer and Playwright documentation describes PDF generation, not preservation of all those items.
For an auditable process, store the PDF alongside whatever provenance and source material your retention policy requires. Validate that page ranges, links, fonts, and essential content appear in the saved document. Do not assume that a visually complete PDF is replayable or contains all source assets.
8. Performance, reliability, and cost
These libraries run browser automation, so your application is responsible for browser installation, process lifecycle, navigation waits, output storage, and handling failures. The research sources provide no measured speed comparison, so there is no supported universal performance winner here.
- Performance: Reusing a browser process for multiple jobs can avoid repeatedly launching a browser, while pages should still be isolated and closed according to your workload and security needs. Benchmark with your target pages, browser version, concurrency, and PDF settings.
- Reliability: Set timeouts, bound retries, close pages and browsers after errors, and make jobs idempotent where possible. A timeout or failed navigation should be reported as a failed capture, not silently treated as a valid archive.
- Cost: Self-hosting has no per-PDF API price in these examples, but it uses compute, memory, storage, and engineering time. Hosted browser infrastructure or a managed capture API adds a service cost; compare it against operating and maintaining your own browsers.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. For a quick screenshot, make one GET request; its API can also return PDFs. See the ScreenshotNeo API documentation for parameters and options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot.
- Bot checks, blank pages, and failed loads are never billed.
- An MCP server lets AI agents use screenshot tools.
- 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
10. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Colors differ from the browser | PDF uses print media and print color adjustment can alter colors. | Check print styles, add -webkit-print-color-adjust: exact, and enable background graphics if needed. |
| Navigation or PDF generation times out | The page has slow resources, long-running requests, or a wait condition that never resolves. | Use a suitable navigation condition, wait for a meaningful selector, and set a bounded timeout. Avoid relying on network idle for pages with persistent traffic. |
| Fonts look wrong or are missing | Fonts may not have loaded or may be unavailable in the runtime environment. | Confirm the page loaded its fonts before capture and install or provide the required fonts in the browser environment. Puppeteer’s PDF guide notes that PDF generation waits for fonts by default. |
| Backgrounds are missing | Background printing is disabled or print CSS removes the backgrounds. | Enable the PDF background option and inspect the page’s print stylesheet. |
| Output is clipped or unexpectedly paginated | Paper size, margins, scaling, or CSS @page rules conflict. |
Set the intended format and margins explicitly; decide whether CSS page size should take precedence; inspect a representative page at the chosen scale. |
| PDF file is absent in Playwright | No output path was supplied, so the API returned bytes rather than writing a file. | Pass path or write the returned PDF buffer yourself. |
| Browser process remains after an error | Cleanup was skipped when an awaited operation threw. | Close the browser in a finally block, as in the examples. |
11. Frequently asked questions
Can I save a page as a PDF without a browser?
A browser-based renderer is useful when the page depends on browser layout and JavaScript. A capture API can handle the browser setup for you; ScreenshotNeo accepts a URL and can return a PDF.
Does the resulting PDF contain the page’s source HTML?
No such preservation is established by the PDF APIs described here. Keep the HTML separately if you need it.
Does choosing Playwright guarantee the same PDF in Firefox, WebKit, and Chromium?
No. Playwright documents support for those engines, but that does not establish cross-engine PDF parity. Validate the engine and version that your workflow will use.
Which should I choose for an existing project?
Prefer the library that already fits the project’s automation and browser requirements. If you need browser engines Playwright documents beyond Chromium, factor that breadth into the choice; test the actual PDF output either way.
