ScreenshotNeo

BlogHTML to image & PDF

How to Convert a Website URL to PDF with the PDFShift API

POST a page URL to PDFShift, check the response, and save the PDF bytes. Includes cURL, Python, Node.js, troubleshooting, and a screenshot API alternative.

By the ScreenshotNeo team4 October 20268 min read

To convert a website URL to PDF with PDFShift, send an HTTP POST request to https://api.pdfshift.io/v3/convert/pdf. Put the page URL in the JSON source property, pass your PDFShift API key in the X-API-Key header, check the HTTP status, and save the successful response body as a .pdf file.

This guide covers the URL workflow in cURL, Python, and Node.js, plus raw HTML input, basic-authenticated pages, failure handling, and practical deployment considerations. PDFShift’s API details below are based on its implementation guides; those guides do not establish current pricing, quotas, a complete error-code list, or a production retry policy.

1. The request and response at a glance

Part Value
Method POST
Endpoint https://api.pdfshift.io/v3/convert/pdf
Authentication API key in the X-API-Key header
Request body JSON with source set to the page URL
Successful response PDF bytes to write to a file

Check the status before saving or serving the response as a finished PDF. A failed HTTP response may contain an error body rather than a PDF. Do not assume every response body is a valid document.

2. Convert a URL with cURL

This command sends the URL as JSON and writes the response body to result.pdf:

curl --fail-with-body \
  --request POST \
  --url https://api.pdfshift.io/v3/convert/pdf \
  --header "X-API-Key: YOUR_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{"source":"https://www.example.com"}' \
  --output result.pdf

Replace YOUR_API_KEY and the example URL. --output saves the response bytes instead of printing them to the terminal. --fail-with-body makes cURL return a failure status for HTTP errors while retaining the response body for diagnosis; if your cURL version does not support that option, use --fail or inspect the HTTP status separately.

3. Convert a URL with Python

Install the HTTP client if needed with python -m pip install requests. Then save this as convert.py and run python convert.py:

import os
import requests

api_key = os.environ["PDFSHIFT_API_KEY"]
source_url = "https://www.example.com"

response = requests.post(
    "https://api.pdfshift.io/v3/convert/pdf",
    headers={"X-API-Key": api_key},
    json={"source": source_url},
    timeout=90,
)
response.raise_for_status()

with open("result.pdf", "wb") as pdf_file:
    pdf_file.write(response.content)

print("Saved result.pdf")

Set the key in the environment before running the script:

export PDFSHIFT_API_KEY="YOUR_API_KEY"
python convert.py

The request uses a timeout so the client does not wait indefinitely. Choose a timeout that fits your application and the pages it must convert. raise_for_status() stops on HTTP errors instead of saving an error response as a PDF. The file is opened in binary mode because PDF output is bytes, not text.

4. Convert a URL with Node.js

This example uses the built-in fetch available in current Node.js releases. Save it as convert.mjs, set PDFSHIFT_API_KEY, and run node convert.mjs.

import { writeFile } from "node:fs/promises";

const apiKey = process.env.PDFSHIFT_API_KEY;
if (!apiKey) {
  throw new Error("Set the PDFSHIFT_API_KEY environment variable");
}

const response = await fetch("https://api.pdfshift.io/v3/convert/pdf", {
  method: "POST",
  headers: {
    "X-API-Key": apiKey,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ source: "https://www.example.com" }),
  signal: AbortSignal.timeout(90_000),
});

if (!response.ok) {
  const detail = await response.text();
  throw new Error(`PDFShift returned HTTP ${response.status}: ${detail}`);
}

const pdfBytes = Buffer.from(await response.arrayBuffer());
await writeFile("result.pdf", pdfBytes);
console.log("Saved result.pdf");

The important Node.js detail is to read the successful response as binary data using arrayBuffer(), then write those bytes to disk. Do not use response.text() for the PDF itself.

5. Use raw HTML when you already have the page content

PDFShift also documents supplying raw HTML as the source. Its raw-HTML guide says this avoids the initial network request to fetch HTML from a URL, can be used for documents that are not publicly accessible, and may reduce conversion duration. It also says inline styles and JavaScript can reduce conversion duration further. These are the vendor’s stated advantages, not independent timing measurements.

Use this approach when your application already has the HTML document and its dependencies can be included appropriately. The request shape remains a JSON source; the value is HTML content instead of a page URL. The exact handling of relative asset paths and additional conversion options should be checked in the current PDFShift documentation for your use case.

curl --fail-with-body \
  --request POST \
  --url https://api.pdfshift.io/v3/convert/pdf \
  --header "X-API-Key: YOUR_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{"source":"<html><body><h1>Invoice</h1><p>Ready to print.</p></body></html>"}' \
  --output result.pdf

For generated content, escape the HTML correctly when constructing JSON; using a JSON serializer in application code is safer than concatenating a request body by hand.

6. Convert a page protected by HTTP Basic Authentication

PDFShift’s PHP guide for secured pages documents an auth object with a username and password for Basic Authentication. This is a specific documented method; do not assume it covers OAuth, form-based sign-in, session cookies, or other authentication schemes.

{
  "source": "https://www.example.com/private-report",
  "auth": {
    "username": "YOUR_USERNAME",
    "password": "YOUR_PASSWORD"
  }
}

Use the documented field shape with the API request, and keep both the PDFShift API key and page credentials out of source control, URLs, and request logs. Confirm that Basic Authentication is the scheme protecting the page before relying on this option.

7. Handle responses and failures safely

  1. Send the POST request. Set the API key header and JSON content type, and provide the source.
  2. Check the HTTP status. Treat a non-success response as a failed conversion and retain its status and diagnostic body for debugging.
  3. Save only successful response bytes. Write in binary mode or use a byte buffer.
  4. Verify the output in your application. If you will expose the file to users, confirm that the saved result can be opened as a PDF before marking the job complete.
  5. Record useful context. Log a request or job identifier from your own system, the source host, elapsed time, and the failure status. Avoid logging secrets or sensitive page contents.

The official examples demonstrate status checks: Python uses raise_for_status(), and the PHP example checks for HTTP 200 before saving. They do not provide a complete retry policy or comprehensive error-code reference. In production, choose retry behavior based on the failure and your application’s deadline; do not blindly retry every error.

8. Troubleshooting

Symptom Likely cause What to do
Authentication or authorization error The API key is missing, invalid, or sent under the wrong header. Send the key as X-API-Key; check the configured secret and avoid adding it to the URL.
Request rejected as invalid The body is not valid JSON, or the source property is missing or malformed. Set Content-Type: application/json, serialize JSON with a library, and verify the source value.
Output file contains an error message or is not a PDF The application saved an unsuccessful response body as if it were a document. Check the HTTP status before writing the file. Inspect the response body as diagnostic text only when the request failed.
Client timeout The request exceeded the client’s configured time limit, possibly while the page or its resources are being rendered. Set a suitable client timeout, investigate whether the source page is reachable and completes rendering, and avoid unlimited waits. The dossier does not establish PDFShift’s server-side timeout behavior.
Page content is missing from the PDF The page may depend on content or assets that are not available to the conversion request, or the page may require authentication. Check the page from the conversion context. For a Basic Authentication page, consult the documented auth option. If you control the content, consider sending raw HTML with appropriate embedded or accessible assets.
Private page still cannot be converted The page may use an authentication scheme other than HTTP Basic Authentication. The sourced secured-page example only documents Basic Authentication. Verify the current PDFShift documentation for any other supported access method; do not assume a browser login session is shared.
PDF appears corrupted after saving The client may have decoded the response as text, or the write path may have changed the bytes. Handle the response as bytes: Python binary file mode, Node.js arrayBuffer() and buffer, or cURL --output.

9. Performance, reliability, and cost considerations

Performance

For a URL source, PDFShift must fetch the page as part of conversion. Its raw-HTML guide says supplying HTML directly removes that initial fetch and can reduce conversion duration; it also recommends embedding styles and JavaScript to reduce network requests. No independent benchmark or guaranteed conversion time is established by the research used for this guide.

Measure your own representative pages before setting client timeouts or job deadlines. Pages with different content and resource dependencies may take different amounts of time. Keep the client timeout finite and make it configurable for the workload.

Reliability

  • Check HTTP status before treating a response as a completed PDF.
  • Keep the API key in a secret store or environment configuration, not in checked-in code or logs.
  • For asynchronous application workflows, track a job as pending until the request succeeds and the bytes have been saved.
  • Use retries selectively. The cited guides do not establish which failures are transient, retry limits, or idempotency behavior, so consult current vendor documentation before implementing automatic retries.

Cost

The implementation guides in the research dossier do not establish current PDFShift pricing, quotas, or billing behavior. Check PDFShift’s current account and pricing information before estimating costs. Avoid assuming that a retry, failed conversion, or particular input mode has a specific charge without confirming the applicable terms.

10. Or skip the browser setup

If you need a screenshot or PDF from a URL without managing browser capture infrastructure, ScreenshotNeo offers a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.

11. Frequently asked questions

Does this endpoint return a PDF URL or the PDF file?

The documented examples treat the successful response body as the PDF bytes and save those bytes to a file.

Can I use a page URL that is not public?

The raw-HTML guide says you can submit HTML for documents that are not publicly accessible. For a URL protected with HTTP Basic Authentication, the secured-page example documents an auth object. Other login mechanisms are not established by the cited material.

Which request library should I choose?

Use the HTTP client already maintained by your application. The core requirements are the same: POST JSON to the v3 endpoint, send the API key in X-API-Key, check the status, and preserve the response as bytes.

Does this guide establish PDFShift’s current limits or price?

No. The implementation sources surfaced for this guide do not establish current pricing, quotas, conversion limits, or a complete error and retry policy. Check the current vendor information for those details.