ScreenshotNeo

BlogHTML to image & PDF

How to Convert Any URL to PDF With an API

Convert reachable web pages to PDF with hosted APIs, cURL, Python, Node.js, browser automation, and practical fixes for dynamic pages.

By the ScreenshotNeo team1 October 20268 min read

Yes. A URL-to-PDF API accepts a web address, renders the page in a browser, and returns PDF bytes. Your application saves those bytes to disk or streams them to storage or a user. The page must be reachable from the rendering service, and authentication, JavaScript, print CSS, bot checks, and dynamic content can affect the result.

Choose an approach

Approach Best for Trade-offs
Hosted URL-to-PDF API Production conversion without operating browsers Provider-specific authentication, request fields, reachability limits, and usage costs
Self-managed Playwright or Puppeteer Full navigation and rendering control You maintain browser binaries, workers, memory, updates, and rendering regressions
HTML-to-PDF API Private or local content that cannot be fetched by a provider You must send HTML, assets, cookies, or an uploaded file

For a hosted service, PDFCrowd documents a form-encoded HTTP API with HTTP Basic authentication. Browserless documents a JSON /pdf endpoint that accepts either a url or html field and returns application/pdf. Their request formats are different, so copy the selected provider’s exact encoding.

What “any URL” means in practice

  • The renderer must be able to reach the URL over the network. A provider cannot fetch your laptop’s localhost address.
  • Private pages need a supported authentication method such as headers, cookies, or an HTML upload.
  • Client-rendered pages may need a wait condition so data appears before PDF generation.
  • Print styles can change layout. In Playwright, CSS @page size takes priority over width, height, or format options.
  • Some sites restrict automated browsers, require a challenge, or block cloud IP ranges.

PDFCrowd: URL to PDF with cURL

PDFCrowd’s documented endpoint is POST https://api.pdfcrowd.com/convert/24.04/. It uses HTTP Basic authentication and form fields; JSON request bodies are not supported. A successful response is PDF bytes with status 200 OK.

curl -u 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
  -F 'url=https://example.com/article' \
  -F 'content_viewport_width=balanced' \
  'https://api.pdfcrowd.com/convert/24.04/' \
  -o article.pdf

The content_viewport_width=balanced field is explicit in the documented example; do not assume it is the service default. Use the API version shown in the endpoint and review its documentation before upgrading.

PDFCrowd with Python

import os
import requests

username = os.environ["PDFCROWD_USERNAME"]
api_key = os.environ["PDFCROWD_API_KEY"]

response = requests.post(
    "https://api.pdfcrowd.com/convert/24.04/",
    auth=(username, api_key),
    files={
        "url": (None, "https://example.com/article"),
        "content_viewport_width": (None, "balanced"),
    },
    timeout=120,
)
response.raise_for_status()
with open("article.pdf", "wb") as pdf_file:
    pdf_file.write(response.content)

PDFCrowd with Node.js

const username = process.env.PDFCROWD_USERNAME;
const apiKey = process.env.PDFCROWD_API_KEY;
const form = new FormData();
form.append('url', 'https://example.com/article');
form.append('content_viewport_width', 'balanced');

const credentials = Buffer.from(`${username}:${apiKey}`).toString('base64');
const response = await fetch('https://api.pdfcrowd.com/convert/24.04/', {
  method: 'POST',
  headers: { Authorization: `Basic ${credentials}` },
  body: form,
});
if (!response.ok) throw new Error(`PDFCrowd returned ${response.status}`);
const pdfBytes = Buffer.from(await response.arrayBuffer());
await require('node:fs').promises.writeFile('article.pdf', pdfBytes);

Browserless: URL to PDF with cURL

Browserless uses a token query parameter and a JSON body. Send exactly one of url or html.

curl -X POST \
  'https://chrome.browserless.io/pdf?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{
    "url": "https://example.com/article",
    "options": {
      "format": "A4",
      "printBackground": true,
      "landscape": false,
      "margin": {"top": "20px", "right": "20px", "bottom": "20px", "left": "20px"}
    }
  }' \
  -o article.pdf

Browserless documents controls for paper format, margins, landscape orientation, page ranges, headers and footers, and background printing. Its PDF endpoint uses Puppeteer under the hood.

Browserless with Python

import requests

response = requests.post(
    "https://chrome.browserless.io/pdf",
    params={"token": "YOUR_TOKEN"},
    json={
        "url": "https://example.com/article",
        "options": {
            "format": "A4",
            "printBackground": True,
            "landscape": False,
            "margin": {"top": "20px", "right": "20px", "bottom": "20px", "left": "20px"},
        },
    },
    timeout=120,
)
response.raise_for_status()
with open("article.pdf", "wb") as pdf_file:
    pdf_file.write(response.content)

Browserless with Node.js

const response = await fetch(
  'https://chrome.browserless.io/pdf?token=YOUR_TOKEN',
  {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      url: 'https://example.com/article',
      options: {
        format: 'A4',
        printBackground: true,
        landscape: false,
        margin: { top: '20px', right: '20px', bottom: '20px', left: '20px' },
      },
    }),
  }
);
if (!response.ok) throw new Error(`Browserless returned ${response.status}`);
const pdfBytes = Buffer.from(await response.arrayBuffer());
await require('node:fs').promises.writeFile('article.pdf', pdfBytes);

Send HTML instead of a URL

Use an HTML or file input when the source is private, local, or assembled by your application. Browserless accepts html instead of url. PDFCrowd documents a text field and file uploads. Include stylesheets and images in forms the renderer can reach, or inline critical assets.

curl -X POST \
  'https://chrome.browserless.io/pdf?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{"html":"<html><body><h1>Invoice</h1></body></html>"}' \
  -o invoice.pdf

Self-managed browser automation

Playwright or Puppeteer is useful when you need to log in, click through a workflow, inspect the DOM, or apply custom navigation logic. The operational cost is yours: package and update the browser, manage workers and memory, and investigate rendering changes. Playwright also notes that headless mode does not support navigation to a PDF document; generate the PDF from the page instead.

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
await page.goto('https://example.com/article', { waitUntil: 'networkidle' });
await page.pdf({
  path: 'article.pdf',
  format: 'A4',
  printBackground: true,
  margin: { top: '20px', right: '20px', bottom: '20px', left: '20px' },
});
await browser.close();

If the page defines CSS @page rules, validate the result because those rules take priority over several API sizing options.

Rendering options to plan for

Need Typical control Validation check
Paper and orientation A4, Letter, custom size, landscape Check page breaks and clipped content
Margins Top, right, bottom, left values Confirm headers and tables do not overlap
Backgrounds Print background graphics Compare colors and contrast with the browser
Dynamic data Wait for a selector, delay, or network idle Assert that a known element contains data
Headers and footers Provider-specific templates or options Render several pages to verify page numbers
Page ranges For example, selected pages only Confirm numbering and blank-page behavior

Reliability, performance, and cost

  • Set a client timeout longer than the provider’s normal render time and retry only transient network or server errors.
  • Use an idempotency key or your own job identifier when a retry could create duplicate records.
  • Store the response as binary bytes; do not decode it as text.
  • Limit concurrency to the provider’s documented quota and your own memory budget.
  • Cache identical conversions when content freshness allows it. Include the URL, relevant options, and an application version in the cache key.
  • Measure page load, rendering, transfer, and storage separately. The research dossier does not establish provider latency, throughput, uptime, retention, or pricing, so verify those terms directly before choosing a service.
  • For self-hosting, budget for browser downloads, worker processes, memory spikes on long pages, and regression checks after browser updates.

Troubleshooting

401 or 403 authentication errors

Check the credential type and location. PDFCrowd expects Basic authentication with username and API key. Browserless expects its token query parameter. Do not send a JSON body to PDFCrowd’s documented endpoint.

The PDF contains an error page or is blank

Open the URL from the renderer’s network perspective, verify redirects and TLS, and confirm the page does not require your local network. For private content, send HTML or use supported headers and cookies. Add a wait condition for client-rendered content.

Localhost or intranet URLs fail

A hosted renderer cannot reach your machine’s localhost. Upload the HTML, expose a secured reachable endpoint, or run the browser inside your network.

Content is missing

Wait for a selector or network idle, increase a deliberate delay, and check lazy-loaded images. Ensure scripts and assets are not blocked by authentication or origin policy.

Layout ignores the requested paper size

Inspect the page’s @page CSS. In Playwright, it takes priority over width, height, or format settings. Remove or override conflicting print CSS when you control the page.

Images or fonts do not appear

Confirm every asset URL is reachable from the renderer, uses valid HTTPS, and does not require browser-only credentials. Inline critical assets for private or short-lived documents.

Requests time out

Find slow third-party scripts, infinite loading indicators, and pages that wait for user interaction. Block unnecessary resources where your provider supports it, or generate from a prepared HTML snapshot.

PDF bytes are corrupted

Write the raw response body to a binary file. Avoid JSON parsing, string conversion, or middleware that changes the response encoding. Log status and content type separately from the body.

Or skip the browser setup

ScreenshotNeo provides a website capture API with PDF output and the rendering controls needed for production pages. It accepts the URL in one GET request; see the ScreenshotNeo API documentation for PDF paper size, margins, landscape mode, and page-range options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can an API convert a page behind a login?

Only if the service supports the required authentication headers, cookies, or HTML input. Otherwise run the renderer where it can access the application.

Should I send a URL or HTML?

Send a URL for public pages. Send HTML or upload a file for localhost, intranet, private, or pre-rendered content.

Why does the PDF differ from the browser tab?

PDF generation uses print media rules, page-size CSS, and a separate rendering environment. Compare print CSS, fonts, asset reachability, and wait conditions.

Is a hosted API better than Playwright?

A hosted API reduces browser operations. Playwright gives deeper control. Choose based on whether operational ownership or navigation control is the bigger requirement.

How should I validate conversion quality?

Use representative pages with long articles, tables, lazy images, custom fonts, print CSS, redirects, and authenticated data. Compare page count, text presence, images, margins, and file validity.