How to Convert a URL to PDF with the DocRaptor API
Create a PDF from a webpage URL with DocRaptor: authenticate, send the request, save the binary response, and troubleshoot JavaScript and fetch errors.
To convert a URL to PDF with DocRaptor, send a JSON POST request to https://api.docraptor.com/docs with type: "pdf" and document_url set to the page to fetch. Authenticate with your API key as the HTTP Basic Auth username and a blank password. For a synchronous request, save the successful response body as binary PDF bytes.
1. Make a PDF with cURL
Set the API key in your shell environment so it is not written into the command or source file. This example uses test mode, which produces a watermarked PDF while you check the output.
export DOCRAPTOR_API_KEY="your_api_key"
curl --user "$DOCRAPTOR_API_KEY:" \
--fail --silent --show-error \
--header "Content-Type: application/json" \
--data '{"test":true,"type":"pdf","document_url":"https://example.com/page"}' \
https://api.docraptor.com/docs \
--output page.pdf
When the PDF is ready for use, remove "test":true. Test mode is for development and its PDFs are watermarked. Keep the live key on a server; putting it in browser code exposes it to visitors.
2. Python: request and save the PDF bytes
Install the HTTP client with python -m pip install requests. This script sends JSON using Basic Auth and writes the successful binary response to a file. It checks the status before saving so an XML error response is not mistaken for a PDF.
import os
import requests
api_key = os.environ["DOCRAPTOR_API_KEY"]
payload = {
"type": "pdf",
"document_url": "https://example.com/page",
"test": True,
}
response = requests.post(
"https://api.docraptor.com/docs",
json=payload,
auth=(api_key, ""),
timeout=180,
)
response.raise_for_status()
with open("page.pdf", "wb") as pdf_file:
pdf_file.write(response.content)
print(f"Saved page.pdf ({len(response.content)} bytes)")
Use a timeout appropriate to your application and expected document size. Do not return the API key to a browser or commit it to source control. For an unwatermarked production document, omit test or set it to false.
3. Node.js: request and save the PDF bytes
Use Node.js 18 or later for the built-in fetch and Buffer APIs. The response is read as an ArrayBuffer, not parsed as JSON.
import { writeFile } from "node:fs/promises";
const apiKey = process.env.DOCRAPTOR_API_KEY;
if (!apiKey) throw new Error("Set DOCRAPTOR_API_KEY first");
const credentials = Buffer.from(`${apiKey}:`).toString("base64");
const response = await fetch("https://api.docraptor.com/docs", {
method: "POST",
headers: {
"Authorization": `Basic ${credentials}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
type: "pdf",
document_url: "https://example.com/page",
test: true,
}),
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
const errorBody = await response.text();
throw new Error(`DocRaptor returned HTTP ${response.status}: ${errorBody}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await writeFile("page.pdf", bytes);
console.log(`Saved page.pdf (${bytes.length} bytes)`);
In production, store the key in server-side configuration. The Authorization header above is equivalent to Basic Auth with the API key as username and an empty password.
4. Configure the document request
DocRaptor accepts either a URL or HTML content as the source. For this URL workflow, use document_url. The output type is pdf. The API reference also documents document_type as a name used by older integrations; use the current type field for new requests.
| Setting | What to use | When it matters |
|---|---|---|
document_url |
The complete URL to fetch, including scheme | The page must be reachable by DocRaptor |
type |
pdf |
Selects PDF output |
| Basic Auth | API key as username; blank password | Required to authenticate the request |
test |
true during layout development |
Test output is watermarked; remove for the final PDF |
javascript |
Omit or leave disabled unless the page needs it | Enables client-side rendering for JS-driven content |
JavaScript is disabled by default. Enable javascript: true for pages whose meaningful content is drawn or populated in the browser, such as a client-rendered chart. It can take longer to render. DocRaptor documents a JavaScript completion signal, docraptorJavaScriptFinished(), for delayed or asynchronous page code. The API reference also warns that enabling both available JavaScript engines can evaluate code twice; only use that configuration when the page requires it.
{
"type": "pdf",
"document_url": "https://example.com/report",
"javascript": true,
"test": true
}
5. Handle responses and choose a request mode
A successful synchronous document request returns PDF bytes in the response body. Write those bytes to a file or stream them to the caller, and check the HTTP status before treating the body as a PDF. An error response contains an XML error message, so avoid blindly saving every response as a .pdf.
The API overview describes other response modes: a hosted document response provides a public URL, while an asynchronous request returns a status_id that can be used to retrieve the document later. Choose synchronous binary output when the caller needs the PDF immediately; hosted or asynchronous handling fits workflows that retrieve the document separately. The X-DocRaptor-Num-Pages response header reports the page count for PDFs.
6. Troubleshoot common problems
| Symptom | Likely cause | Fix |
|---|---|---|
| Authentication fails | The key is missing, invalid, or sent in the wrong place. | Use HTTP Basic Auth with the API key as username and a blank password. Confirm the server process has the intended key configured. |
| The saved file is XML or is not a valid PDF | The request failed and the error body was saved as if it were a successful document. | Check the HTTP status first. Read the error response as text or XML and correct the reported request or fetch problem. |
| “Error downloading document content from supplied url” | DocRaptor could not fetch the supplied page, or the URL does not serve the intended content to its request. | Check the URL spelling and scheme, and confirm the page is reachable by the service without a browser-only login or inaccessible network route. |
| PDF is missing charts or page content | The page relies on JavaScript but JavaScript rendering is off by default. | Set javascript: true. If scripts render asynchronously, use the documented docraptorJavaScriptFinished() signal when applicable. |
| PDF has a watermark | The request is in test mode. | Use test mode for iteration, then remove test: true for an unwatermarked output. |
| Content appears twice after enabling JavaScript | Both available JavaScript engines may be evaluating the page code. | Disable the unnecessary engine and enable only the setup required by the page. |
| Client times out while waiting | Document generation or page rendering exceeded the client timeout. | Set a suitable client timeout and consider the documented asynchronous request mode if the caller should not hold an open request. |
7. Production, reliability, and cost considerations
- Protect credentials: make requests from a backend or another controlled server environment. Do not embed a production API key in public website JavaScript.
- Keep output binary: do not decode the response as UTF-8 or attempt JSON parsing on a successful synchronous response. Stream large documents where your HTTP framework supports it.
- Handle failures explicitly: check HTTP status and preserve the error details for diagnosis. Do not report a PDF as complete until the response succeeds and the bytes have been written.
- Use test mode deliberately: it is useful while adjusting layout, but its watermark makes it unsuitable for a final deliverable.
- Choose rendering settings narrowly: JavaScript is off by default for speed; enabling it only for pages that need client-side rendering avoids unnecessary work.
- Budget by actual plan: DocRaptor’s official Python documentation has stated that paid plans start at $15 per month, but pricing can change. Check the current pricing page before estimating production cost.
8. Or skip the browser setup
If your goal is a clean capture of a page rather than a paginated PDF document, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its API documentation lists options including full-page capture, PDF paper size and margins, page ranges, JavaScript, cookies, custom headers, and async jobs.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Start with the free ScreenshotNeo account.
9. FAQ
Can I convert HTML I already have instead of a URL?
Yes. The API accepts document_content as an alternative source input. This guide uses document_url because it covers converting an existing webpage.
Does a synchronous call return a download link?
Normally it returns the PDF itself as binary data. Hosted document requests return a public URL, and asynchronous requests return a status_id.
How can I tell how many pages were generated?
For PDF responses, inspect the X-DocRaptor-Num-Pages response header.
Should I enable JavaScript on every request?
No. It is disabled by default and should be enabled when the page depends on client-side rendering.


