How to Use the DocRaptor API with Python
Generate PDFs and other documents with DocRaptor's Python client. Learn authentication, safe binary output, test mode, async jobs, and error handling.
Use DocRaptor’s official Python client to send HTML content or a source URL to its document API, then save the returned PDF bytes in binary mode. Install the package with pip install --upgrade docraptor, set your API key as the client’s username, and call create_doc. Use test mode while developing; its generated documents are watermarked.
This guide covers HTML-to-PDF first, then URL input, direct REST requests, output formats, asynchronous jobs, errors, and practical rendering considerations. DocRaptor’s API also supports XLS and XLSX output. Refer to the API overview, API reference, and Python guide for current details.
1. Install the Python client
Install or upgrade the official package in the same environment where your application runs:
python -m pip install --upgrade docraptor
For a project, pin the version you have validated in your dependency file. The research materials do not specify a minimum Python version or a package release number, so check the package metadata and your deployment environment if compatibility is uncertain.
2. Generate a PDF from inline HTML
Save this as generate_pdf.py, set DOCRAPTOR_API_KEY in your environment, then run python generate_pdf.py. The example uses test mode, so the output is watermarked.
import os
import sys
import docraptor
api_key = os.environ.get("DOCRAPTOR_API_KEY")
if not api_key:
sys.exit("Set the DOCRAPTOR_API_KEY environment variable first.")
client = docraptor.DocApi()
client.api_client.configuration.username = api_key
request = {
"test": True,
"document_type": "pdf",
"document_content": """
DocRaptor example
Hello from DocRaptor
This PDF was generated from HTML.
""",
}
try:
response = client.create_doc(request)
except docraptor.rest.ApiException as error:
print(f"DocRaptor request failed: HTTP {error.status} {error.reason}", file=sys.stderr)
if error.body:
print(error.body, file=sys.stderr)
raise SystemExit(1)
with open("document.pdf", "wb") as output:
output.write(bytearray(response))
print("Wrote document.pdf")
The client configuration uses the API key as the authentication username, as shown in DocRaptor’s Python example. Keep the key in an environment variable or a secrets manager; do not commit it in source code, expose it in logs, or include it in a URL.
3. Choose HTML content or a source URL
For a self-contained document or generated report, send document_content. For a page DocRaptor can retrieve, send document_url instead. The API reference requires one of these content sources. Ensure a URL is accessible to DocRaptor and that any required authentication or assets are available to the renderer.
request = {
"test": True,
"document_type": "pdf",
"document_url": "https://example.com/report.html",
}
response = client.create_doc(request)
Do not send both fields unless the current API reference explicitly describes the behavior you want. When generating from HTML, include the styles and resource references needed for the layout. When using a URL, check that relative asset paths resolve as expected from that page.
4. Understand the important request options
| Choice | What to send | When it helps |
|---|---|---|
| Document source | document_content or document_url |
Inline content suits generated reports; a URL suits an existing hosted document. |
| Output type | document_type: "pdf", "xls", or "xlsx" |
Choose the format your workflow consumes. This article’s binary file examples are for PDF. |
| Test mode | test: True |
Useful while developing; DocRaptor says test output is watermarked. |
| Request mode | create_doc or create_async_doc |
Use synchronous generation for requests that finish within the documented synchronous window; use async for longer jobs. |
| PDF rendering options | PDF options documented by DocRaptor and Prince | Use when page layout, headers, accessibility tagging, crop marks, or other PDF-specific rendering behavior matters. |
The REST API currently documents type as the field name and retains document_type for compatibility. The official Python walkthrough uses document_type. Follow the current client and API reference when choosing a field name. Many rendering options are Prince-specific and apply to PDF output; consult DocRaptor’s documentation and the Prince documentation for option definitions and version behavior.
5. Save and serve the binary response safely
A successful direct PDF request returns bytes. Write those bytes with wb, as in the complete example. If your application returns the document over HTTP, stream or write the bytes as a PDF response rather than decoding them as UTF-8 text. A text-mode write can corrupt the file.
The API overview says PDF responses include an X-DocRaptor-Num-Pages header. The Python client example returns response data for writing; if your application needs response headers, confirm how the installed client version exposes them. Do not assume headers are present in the returned byte array.
Hosted-document requests may return a public URL, while asynchronous generation returns a status identifier. Those flows differ from the direct binary response shown above; use the corresponding documented response fields and access controls for your use case.
6. Use asynchronous generation for long-running jobs
DocRaptor’s Python guide describes synchronous generation as limited to 60 seconds and asynchronous generation to 10 minutes. These are vendor-stated limits and may change, so check the current guide before building a hard deadline around them. The async workflow starts with create_async_doc, then polls for status or uses a callback URL to learn when the document is ready.
# Illustrative start of the documented asynchronous workflow.
# Check the current Python guide for the exact status and callback fields.
job = client.create_async_doc({
"test": True,
"document_type": "pdf",
"document_content": "<html><body><h1>Long report</h1></body></html>",
})
print(job)
Async jobs are useful when rendering may take longer than a web request should remain open. Store the returned job identifier, check its status according to the current client documentation, and retrieve the completed document only after the service reports it ready. Make callbacks idempotent so a repeated notification cannot create duplicate downstream work.
7. Call the REST API directly when you do not want the client
The endpoint is https://api.docraptor.com/docs. The API accepts a JSON POST. For direct REST use, the documented HTTP Basic Authentication pattern is the API key as username and a blank password. Prefer that over putting credentials in query parameters.
curl --user "YOUR_API_KEY:" \
-H "Content-Type: application/json" \
-d '{"test":true,"document_type":"pdf","document_content":"<html><body><h1>Hello</h1></body></html>"}' \
https://api.docraptor.com/docs \
--output document.pdf
Python with requests can make the same binary request. Install it with python -m pip install requests, then run:
import requests
response = requests.post(
"https://api.docraptor.com/docs",
auth=("YOUR_API_KEY", ""),
json={
"test": True,
"document_type": "pdf",
"document_content": "<html><body><h1>Hello</h1></body></html>",
},
timeout=90,
)
if not response.ok:
raise RuntimeError(f"HTTP {response.status_code}: {response.text}")
with open("document.pdf", "wb") as output:
output.write(response.content)
The 90-second client timeout above is an example client-side setting, not a DocRaptor service guarantee. Handle error responses before writing a file; error bodies may be XML rather than PDF data.
8. Common errors and fixes
| Symptom | Likely cause | What to do |
|---|---|---|
| Authentication failure | Missing, invalid, or incorrectly configured API key | Check the account key and confirm it is assigned to configuration.username. For REST Basic Auth, use the key as username and an empty password. |
| Request rejected for missing document | Neither document_content nor document_url was supplied |
Provide one valid source field and check spelling against the current API reference. |
| Output file is unreadable | Binary output was decoded or written in text mode; alternatively, an error body was saved as a PDF | Use wb or response.content, check the HTTP outcome first, and inspect the API error body. |
| Unexpected watermark | The request used test: True |
Keep test mode for trial generation; use the appropriate production setting for final output. |
| Python exception with an API status | The service rejected or could not complete the request | Catch docraptor.rest.ApiException; record status, reason, and body while excluding keys and sensitive document content. |
| Request times out in the application | Rendering exceeded the caller’s deadline or the network interrupted the request | Use the async workflow for long jobs, set an appropriate client timeout, and make retries deliberate to avoid duplicate work. |
| Layout or assets differ from expectations | URL access, relative assets, CSS, or Pipeline/Prince version affects rendering | Make required assets reachable, inspect the generated PDF, and verify relevant options against the configured Pipeline version. |
For production diagnostics, keep the HTTP status and service response body where safe, but redact API credentials and private document data. The API overview notes that error responses can be XML, so do not parse every failure as JSON.
9. Rendering, performance, reliability, and cost
Rendering fidelity
DocRaptor identifies Prince as its PDF engine. Prince-specific features include mixed layouts, header placements, accessible PDF tagging, and crop marks. Pipeline versions map to Prince and JavaScript versions, so version changes can affect rendering. Validate representative documents when changing the Pipeline version, styles, source URL, or renderer options.
Performance and reliability
Prefer a URL or inline content based on where the authoritative document lives and how its assets are served. Keep long renders out of latency-sensitive request paths by using asynchronous generation when appropriate. For retries, distinguish a transient network problem from a rejected request, and make downstream file storage and callbacks safe to repeat. The documented synchronous and asynchronous windows are service limits, not a promise that every job completes in that time.
Cost
Plan and test entitlements can change. The research snapshot found a displayed paid-plan starting price and free test details, but those values are volatile; consult DocRaptor’s current pricing page and account terms before estimating production costs. Test output being watermarked does not by itself establish what a particular account includes.
10. Or skip the browser setup
If your actual task is capturing a web page as an image or PDF, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a PNG, JPEG, WebP, or PDF. This is a screenshot workflow rather than a general HTML/XML-to-XLS or document-generation workflow.
Install nothing for this HTTP call; replace the placeholder with your API key. See the ScreenshotNeo API documentation for the request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await Bun.write('shot.webp', res);
For Node.js versions without Bun, write the response bytes with Node’s filesystem API:
import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
- Cookie banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000; every feature is on every plan.
Create a free ScreenshotNeo account to start with 1,000 screenshots per month and no card.
11. Frequently asked questions
Can I use DocRaptor for XLS or XLSX from Python?
Yes. The API reference lists PDF, XLS, and XLSX document types. Check the current format-specific requirements and consume the response using the file type your application expects.
Does test mode make production PDFs?
Test mode is for trial generation and its output is watermarked. Use the production setting for final documents and confirm account terms before deployment.
Should I choose DocRaptor or a screenshot API?
Choose DocRaptor when you need document conversion from HTML/XML or a URL, including spreadsheet formats. Choose a screenshot API when you need a rendered website image or capture-oriented PDF. ScreenshotNeo is one such option, with cleanup for consent banners and widgets and billing that excludes failed or blank captures.
Where do I find the exact async polling fields?
Use the current DocRaptor Python guide and API reference for the installed client version; the async response and callback flow should follow those documented fields.


