How to Capture a Webpage as a PDF with Print CSS Using an API
Generate a paginated PDF from a webpage with print CSS using Puppeteer or Playwright. Includes runnable code, print styling, troubleshooting, and an API option.
To capture a webpage as a PDF with print CSS, render it in a browser automation runtime and call its PDF method. Puppeteer’s page.pdf() and Playwright’s page.pdf() use print media by default, so the browser applies @media print and @page rules. Select screen media before generating the PDF if you want screen styling instead. Both methods return PDF data you can save or serve as application/pdf. Puppeteer API reference · Playwright API reference.
1. Add print CSS to control the document
Print CSS lets a page use a document-friendly layout: hide navigation and interactive controls, set paper size and margins, and reduce awkward page breaks. Add these rules to the page you control:
@media print {
.site-navigation,
.cookie-banner,
.print-button {
display: none !important;
}
article {
max-width: none;
color: #111;
}
h1, h2, h3 {
break-after: avoid;
}
figure, table, blockquote {
break-inside: avoid;
}
}
@page {
size: A4;
margin: 16mm;
}
These are starting rules, not a guarantee of identical pagination across all content. Check long tables, large figures, overflow, and page breaks in the generated PDF. If you do not control the target page, you can sometimes inject CSS through the browser API; the examples below show Puppeteer’s page.addStyleTag().
2. Generate a PDF with Puppeteer
Install Puppeteer in a Node.js project, save this as capture-pdf.mjs, then run it with a URL argument. The script writes page.pdf in the current directory.
npm install puppeteer
// capture-pdf.mjs
import puppeteer from 'puppeteer';
const url = process.argv[2];
if (!url) throw new Error('Usage: node capture-pdf.mjs https://example.com');
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle2', timeout: 60000 });
// Optional: apply print rules when you cannot edit the source stylesheet.
await page.addStyleTag({ content: `
@media print {
.site-navigation, .cookie-banner, .print-button { display: none !important; }
h1, h2, h3 { break-after: avoid; }
figure, table, blockquote { break-inside: avoid; }
}
@page { size: A4; margin: 16mm; }
` });
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', pdf));
} finally {
await browser.close();
}
node capture-pdf.mjs https://example.com
Page.pdf() returns PDF bytes. Puppeteer’s PDF generation waits for fonts by default, according to its guide. Puppeteer PDF generation guide.
Useful Puppeteer PDF options
format: named paper size such asA4orLetter. You can instead supplywidthandheight.printBackground: include background graphics and colors; without it, some backgrounds may be omitted.preferCSSPageSize: prefer the page size declared in CSS@pageover scaling the content to the API paper size.landscape: use landscape orientation when the document is wider than it is tall.margin: set top, right, bottom, and left margins if CSS does not define them.pageRanges: select pages to output, for example'1-3, 5'.scale: adjust content scale when fitting is needed; inspect text size and clipping after changing it.displayHeaderFooter,headerTemplate, andfooterTemplate: add print headers or footers where needed.
Use either CSS page sizing or API sizing deliberately. Conflicting sizes and margins can produce unexpected scaling or whitespace. Puppeteer documents the supported options and print behavior in its Page.pdf() reference.
3. Generate a PDF with Playwright
Playwright’s API follows the same basic flow. Install the package and browser, then save this as capture-playwright.mjs. It creates page.pdf.
npm install playwright
npx playwright install chromium
// capture-playwright.mjs
import { chromium } from 'playwright';
const url = process.argv[2];
if (!url) throw new Error('Usage: node capture-playwright.mjs https://example.com');
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true
});
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', pdf));
} finally {
await browser.close();
}
node capture-playwright.mjs https://example.com
page.pdf() returns a PDF buffer. Playwright also uses print CSS by default and documents PDF options in its Page API reference.
4. Choose print or screen media
For a print stylesheet, leave the default print media selection in place. To generate the PDF with screen styles, explicitly select screen media before calling the PDF method:
// Puppeteer
await page.emulateMediaType('screen');
const pdf = await page.pdf({ format: 'A4', printBackground: true });
// Playwright
await page.emulateMedia({ media: 'screen' });
const pdf = await page.pdf({ format: 'A4', printBackground: true });
Both APIs describe print media as the default for PDF generation and document selecting screen media first when screen styling is wanted. Print output may adjust colors for printing. If exact colors matter, the documentation points to -webkit-print-color-adjust; for example, apply print-color-adjust: exact in a print rule and still verify the resulting file.
@media print {
.brand-panel {
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
}
5. Wait for the content your PDF needs
A page navigation event does not necessarily mean that every application-specific element is ready. Choose a readiness condition that matches the target page:
waitUntil: 'domcontentloaded'for pages where the needed content is already in the initial document.waitUntil: 'networkidle2'in Puppeteer or'networkidle'in Playwright for pages that settle their network activity. Analytics, polling, or streaming requests can prevent an idle condition.- Wait for a known selector after navigation when client-side rendering fills in the document later.
- For lazy-loaded images or sections, scroll the relevant content into view or trigger the application’s own load behavior before printing.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
await page.waitForSelector('article h1', { timeout: 15000 });
await page.evaluate(() => document.fonts.ready);
const pdf = await page.pdf({ format: 'A4', printBackground: true });
Use a finite timeout and handle failure explicitly in production. A page that never reaches network idle may still be printable once its meaningful content is ready; waiting for a specific selector can be more reliable for that kind of application.
6. cURL, Python, and Node.js through a PDF API
If you want a hosted PDF capture instead of maintaining a browser runtime, ScreenshotNeo accepts a URL in one GET request and can return a PDF. Its API also supports PDF settings such as paper size, margins, landscape, and page ranges. See the ScreenshotNeo API documentation for the request options. For exact print stylesheet behavior, verify the PDF against representative pages and CSS rules.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
Python
import requests
response = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://stripe.com",
"format": "pdf",
},
timeout=90,
)
response.raise_for_status()
with open("page.pdf", "wb") as pdf_file:
pdf_file.write(response.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await import('node:fs/promises').then(async fs =>
fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer()))
);
Keep API keys on a server or in a secret manager; do not expose them in browser code or public source repositories. The Node.js example uses the platform’s built-in fetch and requires a Node version that provides it.
7. Or skip the browser setup
ScreenshotNeo can return a PDF from one API call, so you do not have to install and manage a browser runtime for this request. Its clean-shot flow accepts cookie or consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo and its API docs.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
Sign up free for 1,000 screenshots a month, with no card required.
8. Troubleshooting
| Symptom | Likely cause | What to try |
|---|---|---|
| The PDF uses the screen layout | Screen media was selected before printing, or the page has no print-specific rules. | Remove the screen emulation call or select print media, then inspect the page’s @media print rules. |
| Backgrounds or colors are missing | Background printing is disabled, or print color adjustment changes colors. | Set printBackground: true and use -webkit-print-color-adjust: exact for elements that need exact colors; inspect output. |
| Content is clipped or scaled strangely | CSS @page size and API paper size or margins conflict, or a wide element exceeds the printable area. |
Choose one paper-size source, set margins explicitly, and adjust wide tables or content. Try landscape for genuinely wide documents. |
| Images or fonts are missing | Assets had not loaded, require authentication, or are lazy-loaded outside the viewport. | Wait for a meaningful selector and fonts, ensure the capture session can access the assets, and trigger lazy content loading before PDF creation. |
| Navigation hangs at network idle | Long polling, analytics, or persistent requests keep network activity alive. | Wait for DOM content or a page-specific readiness selector rather than network idle. |
| A heading is stranded at the bottom of a page | The print layout allows a break after the heading. | Use break-after: avoid on headings and break-inside: avoid on figures or blocks where keeping content together makes sense. |
| PDF generation times out | The page is slow, blocked, or waiting on an unsuitable readiness condition. | Check URL reachability and authentication, wait for a specific selector, and use a bounded timeout appropriate to your job. |
| API returns an error or an unexpected artifact | Invalid credentials, malformed URL or parameters, or a target that did not load successfully. | Check the response status and headers, URL-encode the target, verify the API key server-side, and consult the API docs. |
9. Performance, reliability, and cost
Browser-based PDF generation has no single reliable speed or cost figure: output time depends on page behavior, assets, browser startup, and your deployment. Reuse browser processes where your service architecture permits, but isolate pages and close them after each job. Bound navigation and rendering time, record failures, and retry only transient failures. For variable third-party pages, inspect output samples and monitor missing content rather than treating successful navigation as proof of a good PDF.
A self-managed browser shifts runtime, memory, concurrency, and maintenance costs to your application. A hosted API avoids operating that browser yourself, but has request and plan limits. ScreenshotNeo lists Free at 1,000 shots monthly with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. Its billing rules exclude bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits; check the response’s verdict and billed headers when accounting for calls.
10. FAQ
Does PDF generation use print CSS automatically?
Yes. Puppeteer and Playwright document print media as the default for their PDF methods.
Can I create a PDF from screen styles instead?
Yes. Select screen media before calling the PDF method.
Will every webpage paginate the same way?
No. Pagination depends on the page’s content, CSS, assets, paper size, margins, and browser rendering. Inspect representative PDFs and tune the print rules.
Can an API create a PDF without Puppeteer or Playwright in my application?
Yes. A hosted capture API such as ScreenshotNeo can accept a URL and return a PDF, avoiding local browser setup for the request.


