ScreenshotNeo

BlogHTML to image & PDF

How to Capture a Webpage as a PDF with Urlbox

Capture a webpage as a PDF with Urlbox using its CLI or API. Choose page layout, paper size, and print styles, then save the result reliably.

By the ScreenshotNeo team4 October 20268 min read

To capture a publicly accessible webpage as a PDF with Urlbox, set format to pdf in an API request, or use the CLI command urlbox pdf https://example.com --output page.pdf. The API returns a temporary render URL that you can download; the CLI writes the PDF to a local file. Urlbox’s API defaults to a multipage PDF, while the CLI’s pdf shortcut enables full-page capture by default.

1. Choose the Urlbox workflow

Use the CLI for a one-off capture from a terminal. Use the API when your application needs to generate PDFs, handle results in code, or process jobs asynchronously. Both methods need content the rendering service can access: a webpage URL must be publicly accessible. The sources for this guide do not establish a procedure for rendering private or internal pages.

Workflow Best for Output handling
CLI Manual or scripted terminal captures Writes to the output path you specify
Synchronous API Application integrations that can wait for rendering Returns a temporary render URL to download
Asynchronous API Long-running or queued capture workflows Returns a render ID and status URL; polling or webhooks can track completion

2. Capture a PDF with the Urlbox CLI

Install and authenticate the Urlbox CLI using its CLI quickstart, then run:

urlbox pdf https://example.com --output page.pdf

The pdf command is shorthand for PDF rendering with full-page capture enabled. If you want a conventional multipage document, use the CLI’s render options to override the full-page setting as documented for your CLI version. Check the resulting PDF: a single long page and a paginated document behave differently in viewers and when printed.

3. Capture and download a PDF with the API

The synchronous endpoint is POST https://api.urlbox.com/v1/render/sync. Authenticate with a bearer token containing your project secret. The minimal render body is:

{
  "url": "https://example.com",
  "format": "pdf"
}

The response includes a temporary renderUrl. Download the PDF from that URL and save it to storage you control if you need to keep it. Urlbox documents a 30-day expiry for synchronous render URLs.

cURL

Replace YOUR_PROJECT_SECRET with your project secret. This example reads the render URL from the JSON response, then downloads the PDF:

curl --fail-with-body -sS \
  -X POST "https://api.urlbox.com/v1/render/sync" \
  -H "Authorization: Bearer YOUR_PROJECT_SECRET" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","format":"pdf"}' \
  -o render.json

RENDER_URL=$(python -c 'import json; print(json.load(open("render.json"))["renderUrl"])')
curl --fail-with-body -L "$RENDER_URL" -o page.pdf

The shell snippet uses Python only to extract the JSON field. If your API response shape differs for your account or API version, inspect render.json and adjust the field access.

Python

import requests

API_URL = "https://api.urlbox.com/v1/render/sync"
PROJECT_SECRET = "YOUR_PROJECT_SECRET"

response = requests.post(
    API_URL,
    headers={"Authorization": f"Bearer {PROJECT_SECRET}"},
    json={"url": "https://example.com", "format": "pdf"},
    timeout=120,
)
response.raise_for_status()
render_url = response.json()["renderUrl"]

pdf_response = requests.get(render_url, timeout=120)
pdf_response.raise_for_status()
with open("page.pdf", "wb") as pdf_file:
    pdf_file.write(pdf_response.content)

Install the dependency with python -m pip install requests. Keep the project secret in an environment variable or secret manager in production rather than committing it to source control.

Node.js

const API_URL = 'https://api.urlbox.com/v1/render/sync';
const projectSecret = process.env.URLBOX_PROJECT_SECRET;
if (!projectSecret) throw new Error('Set URLBOX_PROJECT_SECRET');

const response = await fetch(API_URL, {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${projectSecret}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ url: 'https://example.com', format: 'pdf' }),
  signal: AbortSignal.timeout(120_000),
});
if (!response.ok) throw new Error(`Urlbox render failed: ${response.status} ${await response.text()}`);
const { renderUrl } = await response.json();

const pdfResponse = await fetch(renderUrl, { signal: AbortSignal.timeout(120_000) });
if (!pdfResponse.ok) throw new Error(`PDF download failed: ${pdfResponse.status}`);
const pdfBytes = Buffer.from(await pdfResponse.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', pdfBytes));

This example uses the built-in Fetch API available in current Node.js releases. Set the secret before running it, for example with URLBOX_PROJECT_SECRET in your process environment.

4. Choose PDF layout and page content

Urlbox’s PDF documentation and render options reference describe the settings that most affect the output. Add them to the request body alongside url and format.

Need Option What to expect
Normal paginated document Default behavior The PDF is split into multiple pages.
One long page full_page: true Attempts to put the entire webpage on one PDF page. Very long pages can be awkward to view or print.
Only certain PDF pages pdf_page_range Use a range such as 1,3-5.
Standard paper pdf_page_size Choose a supported standard size.
Custom paper dimensions pdf_page_width and pdf_page_height Set the page dimensions for the output you need.
Use the page’s screen design media: "screen" PDF rendering uses print styles by default; screen styles can preserve content that a print stylesheet changes or omits.
Headers and footers PDF header options Templates can include values such as the date, title, URL, page number, and total page count.
Consent dialog in the capture click_accept and hide_cookie_banners Consider these options for pages where a banner obscures the document; inspect the result to confirm the page content is correct.

For example, a body requesting screen styling and a custom page range can be structured like this; choose supported values and dimensions for your document:

{
  "url": "https://example.com",
  "format": "pdf",
  "media": "screen",
  "pdf_page_range": "1,3-5",
  "pdf_page_size": "A4"
}

Use full_page when a single continuous sheet is explicitly useful, such as for a long reference page. Use ordinary pagination for documents intended to be read or printed as pages. A site’s print stylesheet may intentionally remove navigation, backgrounds, or interactive elements; compare screen styling if the print result misses required content.

5. Handle rendering in an application

A synchronous request waits for rendering and then returns a temporary URL. The Urlbox API reference says long-running synchronous requests can return a temporary redirect after 95 seconds. For a capture that may take longer, use the asynchronous endpoint: it returns a render ID and status URL, which your application can poll; the API also describes webhooks as a completion mechanism.

Plan the download and retention step explicitly. The synchronous render URL expires after 30 days, so download the file into your own storage if it must be retained longer. The API reference also says repeated render-endpoint requests are not cached or deduplicated. Avoid blindly retrying a request after an ambiguous network timeout, because a retry may trigger another render. Track your own job identifier and completion state.

6. Troubleshooting

Symptom Likely cause What to do
Authentication error The bearer token is missing, malformed, or not the project secret. Send Authorization: Bearer ..., verify the secret, and keep it out of client-side code.
The URL cannot be rendered The URL is not publicly accessible to Urlbox. Check access from outside your network. This documented workflow requires a publicly accessible URL; the cited sources do not specify a private-page setup.
Response is JSON, not a PDF The synchronous endpoint response provides a render URL rather than the PDF bytes directly. Read renderUrl from the JSON and make a second request to download the file.
Download fails after rendering The temporary render URL has expired, or the download request failed. Render again if it has expired, download promptly, and check the download response status before saving.
Capture takes too long or request redirects The synchronous render exceeded the documented long-request window. Use the asynchronous workflow and track completion through its status URL or a webhook.
Expected content is missing The page’s print stylesheet omits it, a consent banner obscures it, or the page layout changes for print. Try media: "screen", consider the documented consent-banner options, and inspect the resulting PDF.
Output is a long strip rather than pages Full-page mode was enabled; the CLI pdf shortcut enables it by default. Disable full-page mode for a conventional paginated PDF.
A page range is not applied as expected The range syntax or requested pages do not match the rendered document. Use the documented form such as 1,3-5 and check the page count before selecting ranges.
Saved file is empty or corrupt The response body was saved without checking for HTTP errors, or a JSON error response was mistaken for PDF bytes. Check status codes, parse the render response first, and download from its render URL.

7. Performance, reliability, and cost considerations

  • Rendering time: A synchronous request ties up the caller while the page renders. Use the asynchronous flow when longer jobs should not occupy a request handler.
  • Retries: Since repeated render requests are not cached or deduplicated, implement bounded retries and retain job state so network errors do not cause accidental duplicate captures.
  • Retention: Treat the returned render URL as temporary; download and store the PDF yourself when you need longer retention.
  • Output validation: Check HTTP status and confirm the downloaded response is the expected PDF before storing or exposing it. Review page count, print styles, and consent overlays for important documents.
  • Cost: Urlbox pricing and quotas can change. Check the current Urlbox pricing page before estimating production cost. This guide does not assume a per-render price.

8. Or skip the browser setup

If you need a PDF capture endpoint without maintaining your own rendering workflow, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its endpoint returns a PDF when requested, alongside PNG, JPEG, or WebP screenshots. Cookie banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say whether the page was clean and whether it was billed.

For AI workflows, its MCP server provides take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -d format=pdf \
  -o page.pdf

Sign up for ScreenshotNeo and capture 1,000 shots a month free, with no card.

9. FAQ

Can I capture a page that requires a login?

The Urlbox sources used here establish the public-URL workflow, not a specific private-page authentication procedure. Do not assume a login-protected URL will be accessible through the basic request.

Should I use PDF output or a full-page image?

Use PDF when you need page-oriented output, selectable document content, or page ranges and paper sizing. Full-page capture controls whether the PDF is one long page; it is still a PDF output.

Can I keep the API result URL forever?

No. The synchronous API documentation gives the temporary render URL a 30-day expiry. Download the PDF and retain it in your own storage for longer-term access.

Why does the CLI result differ from the API default?

The CLI’s pdf shortcut enables full-page capture, whereas the documented API PDF default is multipage. Set the full-page option deliberately when matching outputs across workflows.

Sources