How to Convert Web Pages to PDF with an API
Convert a URL to PDF with a browser or hosted API. Learn how to control print styles, page size, output handling, and common failures.
To convert a web page to PDF with an API, render the page in a browser engine and call its PDF-generation method, or send a URL or HTML to a hosted PDF API. Choose print CSS when you want a document layout, screen CSS when you want the page to look like it does in a browser, then set the page size, margins, and output handling your workflow requires.
This guide shows a self-hosted Node.js workflow with Puppeteer, explains the equivalent Playwright approach, and covers hosted conversion APIs. The browser examples use the APIs documented by Puppeteer and Playwright. Check the documentation for your installed version before relying on a particular option.
1. Choose a conversion approach
| Approach | Good fit when | What to account for |
|---|---|---|
| Browser automation library | You need rendering control, access to pages your application can reach, or a browser workflow you can operate yourself. | Your application must run the browser, manage its lifecycle, and handle timeouts and output storage. |
| Hosted PDF API | You want to submit a URL or HTML to a service endpoint rather than run the browser in your application. | Verify its supported inputs, options, response format, authentication, security terms, limits, and pricing in that provider’s documentation. |
For a browser-based implementation, Puppeteer and Playwright both expose a page PDF method. Their documented defaults use print CSS. They also document selecting screen media before PDF generation. A hosted API has its own contract: for example, the Chromium PDF Service API reference has distinct HTML, URL, and file routes. Those routes are an example of one API design, not a promise that every provider accepts all three inputs.
2. Convert a URL to PDF with Puppeteer in Node.js
Install Puppeteer in a Node.js project:
npm install puppeteer
Save this as url-to-pdf.mjs. It navigates to a URL, waits for the page load event, and writes an A4 PDF to disk:
import puppeteer from 'puppeteer';
const url = process.argv[2] ?? 'https://example.com';
const output = process.argv[3] ?? 'page.pdf';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto(url, {
waitUntil: 'load',
timeout: 30_000,
});
await page.pdf({
path: output,
format: 'A4',
printBackground: true,
margin: {
top: '12mm',
right: '12mm',
bottom: '12mm',
left: '12mm',
},
});
console.log(`Saved ${output}`);
} finally {
await browser.close();
}
Run it with node url-to-pdf.mjs https://example.com report.pdf. The finally block closes Chromium even if navigation or PDF creation fails. printBackground includes background graphics where supported by the installed Puppeteer version.
Set print or screen CSS
PDF generation normally uses print styles. That is often right for a report: sites may hide navigation, expand content, or adjust spacing in print CSS. To render screen styles instead, emulate the screen media type before calling page.pdf():
await page.emulateMediaType('screen');
await page.pdf({ path: 'screen-style.pdf', format: 'A4', printBackground: true });
When printed colors look faded or backgrounds are missing, Puppeteer documents using -webkit-print-color-adjust in page CSS to request exact colors:
@media print {
html {
-webkit-print-color-adjust: exact;
}
}
Color adjustment can make a PDF less printer-friendly and may increase ink use. Use it when visual color fidelity matters more than print economy.
3. Configure the PDF output
Choose page dimensions and rendering behavior to match the document’s intended use. Puppeteer and Playwright expose related controls, but option names and availability can depend on library and version.
| Need | What to configure |
|---|---|
| Standard paper | Select a documented format such as A4 or Letter. Confirm the format name supported by your library version. |
| Custom page size | Set width and height with supported units. Playwright documents unit-bearing dimensions and margins; confirm syntax in the chosen library reference. |
| More or less white space | Set top, right, bottom, and left margins. Leave enough room for headers, footers, or printer-safe content. |
| Background graphics | Enable the library’s background-printing option if the PDF should retain CSS backgrounds. |
| Smaller or larger rendering | Use the PDF scale option where available, and inspect the result for clipped text and unexpected page breaks. |
| Page numbering or labels | Puppeteer supports header and footer templates through its PDF options. Check the template requirements for the installed version. |
| Screen appearance | Set screen media before PDF generation if that is the intended design. The PDF call itself otherwise uses print media in the documented defaults. |
| Exact colors | Use the documented print color adjustment CSS when needed and verify the resulting file. |
With Puppeteer, path saves the result to a file; a relative path resolves against the process working directory. Without a path, the documented PDF method returns PDF bytes, which you can pass to storage or an HTTP response. Puppeteer also documents createPDFStream() when a readable stream better fits your output pipeline. See Puppeteer’s PDF options, Page.pdf(), and Page.createPDFStream() for details.
4. Convert a URL with Playwright
Playwright offers the same broad workflow: launch a browser, navigate to a page, and call page.pdf(). Install the package and its browser using the commands in the Playwright installation guide. This runnable example assumes Chromium is installed for Playwright:
import { chromium } from 'playwright';
const browser = await chromium.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'load',
timeout: 30_000,
});
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' },
});
} finally {
await browser.close();
}
For screen CSS, set the page media before generating the PDF:
await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'screen-style.pdf', format: 'A4', printBackground: true });
Consult the current Playwright Page reference for PDF settings, units, and supported behavior.
5. Call a hosted PDF conversion API
A hosted service can accept a URL or HTML and return a PDF response. The request method, route, authentication, option names, and response handling are provider-specific. The Chromium PDF Service API reference, for example, documents separate endpoints for HTML, URL, and file inputs. Use that service’s documentation to construct an actual request; do not assume an endpoint shape from another provider will work.
Before adopting a hosted API, check:
- Which inputs it accepts: a public URL, HTML string, uploaded file, or some subset.
- How it authenticates requests and protects submitted content.
- Whether it supports print or screen media, page dimensions, margins, background graphics, and headers or footers.
- Whether the response is PDF bytes, a download URL, or an asynchronous job result.
- How it reports navigation errors, timeouts, unsupported content, and conversion failures.
- Its current limits, retention policy, security terms, availability commitments, and pricing.
Those details are not consistent across providers, so rely on the specific service’s current API documentation and terms.
6. Return the generated PDF from a Node.js server
If an HTTP endpoint in your application creates the PDF, return the generated bytes with a PDF content type and a download filename. This example shows the response shape after you have a Puppeteer page ready:
const pdf = await page.pdf({ format: 'A4', printBackground: true });
response.writeHead(200, {
'Content-Type': 'application/pdf',
'Content-Disposition': 'attachment; filename="page.pdf"',
'Content-Length': String(pdf.length),
});
response.end(Buffer.from(pdf));
For large outputs or high concurrency, consider stream-based handling where the library and server framework support it. Define a maximum page-generation time, close pages and browsers when requests finish, and avoid returning a success response before the PDF bytes are ready.
7. Handle dynamic pages and unreliable navigation
A page’s load event does not guarantee that every asynchronous widget or client-rendered section is complete. Conversely, waiting for every network request can hang on pages that keep analytics or long polling connections open. Choose a readiness condition based on the page:
- Use a navigation event such as
loadas a simple starting point. - If the content appears later, wait for a selector that identifies the content you need.
- Use a bounded timeout and record the URL and failure phase when navigation or rendering fails.
- For pages you control, expose a reliable ready marker after required data and fonts are loaded.
- Capture a representative PDF during development and check page breaks, missing images, and font fallback.
Only automate pages you are authorized to access. For authenticated content, use an approved session or credentials flow, keep secrets out of logs, and confirm that your browser environment can reach the page. Hosted services may have different access and security constraints; verify them with the provider.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The PDF is blank or mostly empty | The page had not rendered its main content before the PDF call, or the page returned a bot check or error page. | Wait for a content selector or app-ready marker; inspect the rendered page and navigation status before creating the PDF. |
| Styles differ from the browser view | PDF generation is using print CSS, or print rules change layout and hide elements. | Use the print layout intentionally, or emulate screen media before generating the PDF. |
| Backgrounds or colors are missing | Background printing is disabled or print color adjustment changes colors. | Enable background printing and apply the documented print color CSS if exact colors are required. |
| Images, charts, or fonts are missing | Resources had not loaded, a request failed, or the browser environment lacks access to the resource. | Wait for the relevant content, check resource access and console/network errors, and ensure required fonts are available to the browser. |
| Navigation times out | The site is slow, unreachable, or keeps network activity open. | Set a bounded timeout, wait for a more appropriate event or selector, and capture diagnostics. Avoid unbounded waits. |
| Content is clipped or split badly | Page size, margins, scale, or print CSS do not suit the content. | Adjust dimensions, margins, scale, and print rules; inspect long tables and large images across page breaks. |
| PDF colors look muted | Browser print color adjustments are applied. | Use -webkit-print-color-adjust: exact when color fidelity is required, then review the PDF output. |
| The process exhausts memory or stalls under load | Too many pages or browser instances are active, or large pages generate large PDFs. | Limit concurrency, reuse browser processes carefully, close pages after each job, and stream or store output according to its size. |
| The hosted API rejects the request | The route, input type, option, authentication, or payload format does not match that provider’s contract. | Check the provider’s current endpoint reference and error response; do not infer parameter names from another API. |
9. Performance, reliability, and cost
Browser-based PDF generation consumes browser and application resources. Keep concurrency bounded, use finite navigation and generation timeouts, and release pages and browser processes when work completes. Large or long pages can take longer to render and produce larger files. For a service that receives jobs from users, isolate conversion work from latency-sensitive request handling when practical and retain enough diagnostics to identify navigation, rendering, and output failures.
For repeat jobs, caching may avoid rerendering unchanged content, but only when the page and credentials make reuse safe. A hosted API can reduce browser operations in your application, but its operational behavior, data handling, quotas, and pricing must be evaluated from that provider’s current terms. The research-backed references establish browser methods and an example of hosted endpoint types; they do not establish comparative costs, security guarantees, quotas, or service reliability.
10. Or skip the browser setup
ScreenshotNeo is a website screenshot API that can also return PDFs. One GET request takes a URL. See the ScreenshotNeo API documentation for PDF options and other parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('page.pdf', Buffer.from(await res.arrayBuffer())));
With ScreenshotNeo, cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.
11. Frequently asked questions
Should I submit a URL or HTML?
Submit a URL when the target is reachable to the renderer and should be loaded as a page. Submit HTML when your workflow already has the markup and the service supports that input. Check whether the API also needs assets or a base URL to resolve relative links.
Does a PDF preserve the page exactly as it looks on screen?
Not by default. Browser PDF methods use print CSS in their documented default behavior. Select screen media when that is the intended appearance, and check the output because PDF pagination still differs from a scrolling viewport.
Can I generate a PDF without saving a temporary file?
Yes. Puppeteer’s documented PDF method returns bytes when no output path is supplied, and its stream method returns a readable stream. A hosted API may return bytes or another response form; confirm its contract.
How do I choose a PDF API?
Compare accepted inputs, rendering options, output transport, authentication, security terms, limits, and current price. The cited API references do not establish those commercial or operational details for providers generally.


