How to Turn a Website URL into a PDF Using an API
Convert a website URL to a downloadable PDF with an API. Learn the request format, rendering options, dynamic-page limits, error handling, and alternatives.
To turn a website URL into a PDF with an API, send an authenticated HTTP request to a PDF-rendering endpoint with the URL and rendering options. The endpoint returns PDF bytes; save those bytes as a binary .pdf file. For example, Browserless documents a POST to its /pdf endpoint with a JSON body. Keep the API token on your server, check the HTTP status, and do not treat an error response as a PDF.
This guide uses Browserless for the direct PDF API example because the research sources document its request and options. The exact endpoint, authentication, and available fields are provider-specific; consult the current provider schema before shipping.
1. Make a URL-to-PDF request
Set the target page in url, choose print settings under options, and write the response body to a file. Replace the placeholder token with a valid credential stored in a server-side secret or environment variable.
curl -X POST 'https://production-sfo.browserless.io/pdf?token=YOUR_API_TOKEN' \
-H 'Content-Type: application/json' \
-H 'Accept: application/pdf' \
-d '{
"url": "https://example.com/",
"options": {
"format": "A4",
"printBackground": true,
"displayHeaderFooter": true
}
}' \
-o page.pdf
The response is binary PDF data. The -o option saves it without interpreting it as terminal text. Browserless documents the same general request pattern in its PDF API documentation.
Python: check status and save the bytes
import os
import requests
endpoint = "https://production-sfo.browserless.io/pdf"
token = os.environ["BROWSERLESS_TOKEN"]
payload = {
"url": "https://example.com/",
"options": {
"format": "A4",
"printBackground": True,
"displayHeaderFooter": True,
},
}
response = requests.post(
endpoint,
params={"token": token},
json=payload,
headers={"Accept": "application/pdf"},
timeout=90,
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "application/pdf" not in content_type.lower():
raise RuntimeError(f"Expected a PDF, got Content-Type: {content_type!r}")
with open("page.pdf", "wb") as pdf_file:
pdf_file.write(response.content)
Node.js: check status and write the bytes
import { writeFile } from "node:fs/promises";
const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error("Set BROWSERLESS_TOKEN first");
const endpoint = new URL("https://production-sfo.browserless.io/pdf");
endpoint.searchParams.set("token", token);
const response = await fetch(endpoint, {
method: "POST",
headers: {
"Content-Type": "application/json",
"Accept": "application/pdf",
},
body: JSON.stringify({
url: "https://example.com/",
options: {
format: "A4",
printBackground: true,
displayHeaderFooter: true,
},
}),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) {
const detail = await response.text();
throw new Error(`PDF API returned ${response.status}: ${detail}`);
}
const contentType = response.headers.get("content-type") ?? "";
if (!contentType.toLowerCase().includes("application/pdf")) {
throw new Error(`Expected a PDF, got Content-Type: ${contentType}`);
}
await writeFile("page.pdf", Buffer.from(await response.arrayBuffer()));
These examples intentionally test the status and content type before saving. Error responses may contain readable diagnostic text, not PDF bytes. Avoid logging tokens or sending them to browser code, mobile clients, or public repositories.
2. Choose the input and PDF layout
Remote URL or supplied HTML
Use url when the renderer should navigate to a page. Use html when your application already has markup to render. Browserless documents these as alternative inputs and says not to send both fields in the same request. Check the endpoint schema for the accepted HTML field shape and supported options.
Common layout controls
| Setting | When to use it |
|---|---|
| Paper format | Choose a standard size such as A4 or Letter for documents that should print consistently. |
| Custom dimensions | Use when the output needs a nonstandard page size; verify the provider’s units and schema. |
| Margins | Reserve space for readable text, binding, or headers and footers. Check whether the API expects units or strings. |
| Landscape | Use for wide tables, charts, or layouts that would otherwise be clipped or split awkwardly. |
| Page ranges | Generate only needed pages when the endpoint supports ranges; confirm its syntax in the live schema. |
| Background printing | Enable it when background colors or images are part of the intended design. Without it, print output can omit them. |
| Headers and footers | Enable them when page numbers or document context matter. Check the endpoint’s template and margin requirements. |
| Tagged output | Where supported, structural tags can help assistive technology and document parsers, but do not guarantee PDF/UA conformance. |
Browserless also documents navigation and waiting controls, cookies, request interceptors, user-agent configuration, and resource blocking. Those are provider-specific. Use the live schema rather than assuming another service accepts identical names or semantics.
3. Handle print styles and dynamic pages
PDF generation follows print rendering by default. A page can therefore look different in the PDF than in a normal browser window: websites may have dedicated print styles, omit interactive elements, or change page breaks. Puppeteer documents that page.pdf() uses the print CSS media type. If you need screen styling, emulate screen media before creating the PDF in a browser-automation workflow. For exact printed colors, the CSS property -webkit-print-color-adjust can affect color adjustment.
For a simple page that is ready on navigation, a single REST call is convenient. It is a poor fit for a workflow that must log in, click controls, navigate through multiple steps, or inspect page state before printing. Browserless describes its REST calls as stateless, single actions: cookies and browser state do not persist between requests. For interaction, use a browser connection or custom browser function that can wait for the right condition and then call PDF generation.
Do not assume every destination permits automated access. Bot detection and interactive challenges can prevent a render, and endpoint schemas may reject forbidden destinations. Respect the site’s access rules and return failures clearly to the caller instead of silently producing an empty or incorrect document.
4. Know the PDF’s structure and limits
- Paginated output: a normal PDF divides content across pages. Browserless’s documented
/pdfREST endpoint does not create one continuous, very tall page for an entire website. Its documentation points to a custom function workflow for that special layout. - Selectable text: Browserless says its generated PDFs contain selectable text rather than being screenshots. This is useful for search and copying, but page structure still depends on the source page.
- Accessibility tags: Browserless documents a
tagged: trueoption for structural information such as reading order and headings. Tagged output is not certified PDF/UA; validate separately if you have a formal conformance requirement. - Page breaks: long tables, cards, and images may split across pages. Adjust the source’s print CSS where you control it, and inspect representative output at the chosen paper size.
These details are documented in Browserless’s PDF guidance. For production documents, validate output with representative pages and any accessibility or archival checks your use case requires.
5. Troubleshoot common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Malformed request or client error | Invalid JSON, unsupported option, wrong field type, or both URL and HTML supplied. | Validate the JSON and compare each field with the current endpoint schema. Send one input mode. |
| Unauthorized response | Missing, expired, or invalid API token. | Check the server-side secret and the provider’s authentication format. Do not put the token in public client code. |
| Forbidden destination | The provider blocks the destination or the site disallows access. | Check the provider’s destination policy and the site’s access rules. Do not try to bypass an access restriction. |
| Timeout or blank output | The page is slow, requires client-side rendering, or did not reach the expected state before capture. | Use documented wait or navigation controls. For interaction or precise readiness checks, move to a browser session workflow. |
| Missing colors or backgrounds | Print rendering omits backgrounds or applies print-specific styles. | Enable background printing and inspect the page’s print CSS. Use screen media emulation if the desired output is the screen view. |
| Rate limit | Request volume exceeded the endpoint’s allowed rate. | Handle the documented rate-limit response, queue work, and retry with bounded backoff where appropriate. |
| Server error or unavailable service | Temporary provider-side failure. | Record a request identifier if returned, surface the failure, and retry only under a bounded policy. Make jobs safe to retry. |
| Saved file is not a PDF | The code saved an error body or proxy response as if it were PDF bytes. | Check status and content type before writing; inspect the error body on failure. |
The documented error classes include malformed requests, authentication failures, forbidden destinations, timeouts, rate limits, server errors, and service unavailability. Exact status codes and retry guidance can vary by endpoint; follow its current documentation.
6. Production, performance, and cost considerations
- Keep credentials private. Make the request from a backend service. Treat the token as a secret and avoid exposing it in logs, URLs shared with users, or client-side source.
- Set a finite timeout. A page can hang on slow resources or never settle. Choose a timeout appropriate to your job and return a useful failure state.
- Make retries bounded and safe. Retry transient timeouts or server failures selectively, with a limit and backoff. Do not retry invalid input or authentication errors unchanged.
- Control concurrency. Queue large batches and honor provider rate limits. A browser render consumes more work than a small JSON request; avoid launching unbounded jobs from user traffic.
- Measure the output path. Track request outcome, render duration, output size, and failure reason in your own system. The available research does not establish comparable provider speed or reliability benchmarks.
- Budget from current terms. The research sources do not establish current Browserless or DocRaptor pricing. Check the provider’s current plan, usage rules, and overage terms before estimating spend.
- Validate representative pages. Fonts, remote images, page length, print styles, and dynamic content all affect output. Keep sample documents that cover the layouts your application supports.
For a site you control, DocRaptor documents a specialized referrer-based route using /docs/from_site: configure allowed domains, then link to that endpoint from the controlled site. Its documentation says JavaScript and PDF options can be passed as query parameters. This is intended for a controlled site, not arbitrary public URL conversion. See the DocRaptor from-site documentation for the constraints.
7. Or skip the browser setup
If you need a screenshot or visual capture rather than a multi-page, selectable-text PDF, ScreenshotNeo offers a one-request website capture API and an MCP server for AI agents. The PDF endpoint returns a PDF; ScreenshotNeo’s documented call below is a screenshot call, so use it when an image capture fits the job.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options and response details. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. If a clean screenshot is what you need, sign up free for 1,000 screenshots a month, with no card.
8. Frequently asked questions
Can I convert a page that requires a login?
A stateless REST call does not retain a browser session between calls. If the page needs authentication or several interactive steps, use a documented cookie mechanism when appropriate or a browser session workflow that can perform the steps before generating the PDF.
Will the PDF look exactly like the page on screen?
Not automatically. PDF rendering commonly uses print styles, which may intentionally differ from screen styles. Background printing and screen media emulation address different parts of that difference.
Can I make a single-page PDF as tall as the whole website?
Not with the documented Browserless /pdf REST endpoint. Its documentation directs that special case to a custom function workflow.
Does a tagged PDF guarantee accessibility compliance?
No. Tags can help expose structure, but the source markup and final document still need evaluation against the requirements you follow.
Can I use an API to make a PDF from HTML I already generated?
Yes, when the endpoint accepts raw HTML. Send the HTML input instead of a URL and follow the provider’s schema; do not submit both input fields together when the endpoint forbids it.


