ScreenshotNeo

BlogHow-to

How to Use the DocRaptor API with Python

Generate PDFs and other documents with DocRaptor's Python client. Learn authentication, safe binary output, test mode, async jobs, and error handling.

By the ScreenshotNeo team4 October 20269 min read

Use DocRaptor’s official Python client to send HTML content or a source URL to its document API, then save the returned PDF bytes in binary mode. Install the package with pip install --upgrade docraptor, set your API key as the client’s username, and call create_doc. Use test mode while developing; its generated documents are watermarked.

This guide covers HTML-to-PDF first, then URL input, direct REST requests, output formats, asynchronous jobs, errors, and practical rendering considerations. DocRaptor’s API also supports XLS and XLSX output. Refer to the API overview, API reference, and Python guide for current details.

1. Install the Python client

Install or upgrade the official package in the same environment where your application runs:

python -m pip install --upgrade docraptor

For a project, pin the version you have validated in your dependency file. The research materials do not specify a minimum Python version or a package release number, so check the package metadata and your deployment environment if compatibility is uncertain.

2. Generate a PDF from inline HTML

Save this as generate_pdf.py, set DOCRAPTOR_API_KEY in your environment, then run python generate_pdf.py. The example uses test mode, so the output is watermarked.

import os
import sys

import docraptor

api_key = os.environ.get("DOCRAPTOR_API_KEY")
if not api_key:
    sys.exit("Set the DOCRAPTOR_API_KEY environment variable first.")

client = docraptor.DocApi()
client.api_client.configuration.username = api_key

request = {
    "test": True,
    "document_type": "pdf",
    "document_content": """

DocRaptor example

Hello from DocRaptor

This PDF was generated from HTML.

""", } try: response = client.create_doc(request) except docraptor.rest.ApiException as error: print(f"DocRaptor request failed: HTTP {error.status} {error.reason}", file=sys.stderr) if error.body: print(error.body, file=sys.stderr) raise SystemExit(1) with open("document.pdf", "wb") as output: output.write(bytearray(response)) print("Wrote document.pdf")

The client configuration uses the API key as the authentication username, as shown in DocRaptor’s Python example. Keep the key in an environment variable or a secrets manager; do not commit it in source code, expose it in logs, or include it in a URL.

3. Choose HTML content or a source URL

For a self-contained document or generated report, send document_content. For a page DocRaptor can retrieve, send document_url instead. The API reference requires one of these content sources. Ensure a URL is accessible to DocRaptor and that any required authentication or assets are available to the renderer.

request = {
    "test": True,
    "document_type": "pdf",
    "document_url": "https://example.com/report.html",
}
response = client.create_doc(request)

Do not send both fields unless the current API reference explicitly describes the behavior you want. When generating from HTML, include the styles and resource references needed for the layout. When using a URL, check that relative asset paths resolve as expected from that page.

4. Understand the important request options

Choice What to send When it helps
Document source document_content or document_url Inline content suits generated reports; a URL suits an existing hosted document.
Output type document_type: "pdf", "xls", or "xlsx" Choose the format your workflow consumes. This article’s binary file examples are for PDF.
Test mode test: True Useful while developing; DocRaptor says test output is watermarked.
Request mode create_doc or create_async_doc Use synchronous generation for requests that finish within the documented synchronous window; use async for longer jobs.
PDF rendering options PDF options documented by DocRaptor and Prince Use when page layout, headers, accessibility tagging, crop marks, or other PDF-specific rendering behavior matters.

The REST API currently documents type as the field name and retains document_type for compatibility. The official Python walkthrough uses document_type. Follow the current client and API reference when choosing a field name. Many rendering options are Prince-specific and apply to PDF output; consult DocRaptor’s documentation and the Prince documentation for option definitions and version behavior.

5. Save and serve the binary response safely

A successful direct PDF request returns bytes. Write those bytes with wb, as in the complete example. If your application returns the document over HTTP, stream or write the bytes as a PDF response rather than decoding them as UTF-8 text. A text-mode write can corrupt the file.

The API overview says PDF responses include an X-DocRaptor-Num-Pages header. The Python client example returns response data for writing; if your application needs response headers, confirm how the installed client version exposes them. Do not assume headers are present in the returned byte array.

Hosted-document requests may return a public URL, while asynchronous generation returns a status identifier. Those flows differ from the direct binary response shown above; use the corresponding documented response fields and access controls for your use case.

6. Use asynchronous generation for long-running jobs

DocRaptor’s Python guide describes synchronous generation as limited to 60 seconds and asynchronous generation to 10 minutes. These are vendor-stated limits and may change, so check the current guide before building a hard deadline around them. The async workflow starts with create_async_doc, then polls for status or uses a callback URL to learn when the document is ready.

# Illustrative start of the documented asynchronous workflow.
# Check the current Python guide for the exact status and callback fields.
job = client.create_async_doc({
    "test": True,
    "document_type": "pdf",
    "document_content": "<html><body><h1>Long report</h1></body></html>",
})
print(job)

Async jobs are useful when rendering may take longer than a web request should remain open. Store the returned job identifier, check its status according to the current client documentation, and retrieve the completed document only after the service reports it ready. Make callbacks idempotent so a repeated notification cannot create duplicate downstream work.

7. Call the REST API directly when you do not want the client

The endpoint is https://api.docraptor.com/docs. The API accepts a JSON POST. For direct REST use, the documented HTTP Basic Authentication pattern is the API key as username and a blank password. Prefer that over putting credentials in query parameters.

curl --user "YOUR_API_KEY:" \
  -H "Content-Type: application/json" \
  -d '{"test":true,"document_type":"pdf","document_content":"<html><body><h1>Hello</h1></body></html>"}' \
  https://api.docraptor.com/docs \
  --output document.pdf

Python with requests can make the same binary request. Install it with python -m pip install requests, then run:

import requests

response = requests.post(
    "https://api.docraptor.com/docs",
    auth=("YOUR_API_KEY", ""),
    json={
        "test": True,
        "document_type": "pdf",
        "document_content": "<html><body><h1>Hello</h1></body></html>",
    },
    timeout=90,
)

if not response.ok:
    raise RuntimeError(f"HTTP {response.status_code}: {response.text}")

with open("document.pdf", "wb") as output:
    output.write(response.content)

The 90-second client timeout above is an example client-side setting, not a DocRaptor service guarantee. Handle error responses before writing a file; error bodies may be XML rather than PDF data.

8. Common errors and fixes

Symptom Likely cause What to do
Authentication failure Missing, invalid, or incorrectly configured API key Check the account key and confirm it is assigned to configuration.username. For REST Basic Auth, use the key as username and an empty password.
Request rejected for missing document Neither document_content nor document_url was supplied Provide one valid source field and check spelling against the current API reference.
Output file is unreadable Binary output was decoded or written in text mode; alternatively, an error body was saved as a PDF Use wb or response.content, check the HTTP outcome first, and inspect the API error body.
Unexpected watermark The request used test: True Keep test mode for trial generation; use the appropriate production setting for final output.
Python exception with an API status The service rejected or could not complete the request Catch docraptor.rest.ApiException; record status, reason, and body while excluding keys and sensitive document content.
Request times out in the application Rendering exceeded the caller’s deadline or the network interrupted the request Use the async workflow for long jobs, set an appropriate client timeout, and make retries deliberate to avoid duplicate work.
Layout or assets differ from expectations URL access, relative assets, CSS, or Pipeline/Prince version affects rendering Make required assets reachable, inspect the generated PDF, and verify relevant options against the configured Pipeline version.

For production diagnostics, keep the HTTP status and service response body where safe, but redact API credentials and private document data. The API overview notes that error responses can be XML, so do not parse every failure as JSON.

9. Rendering, performance, reliability, and cost

Rendering fidelity

DocRaptor identifies Prince as its PDF engine. Prince-specific features include mixed layouts, header placements, accessible PDF tagging, and crop marks. Pipeline versions map to Prince and JavaScript versions, so version changes can affect rendering. Validate representative documents when changing the Pipeline version, styles, source URL, or renderer options.

Performance and reliability

Prefer a URL or inline content based on where the authoritative document lives and how its assets are served. Keep long renders out of latency-sensitive request paths by using asynchronous generation when appropriate. For retries, distinguish a transient network problem from a rejected request, and make downstream file storage and callbacks safe to repeat. The documented synchronous and asynchronous windows are service limits, not a promise that every job completes in that time.

Cost

Plan and test entitlements can change. The research snapshot found a displayed paid-plan starting price and free test details, but those values are volatile; consult DocRaptor’s current pricing page and account terms before estimating production costs. Test output being watermarked does not by itself establish what a particular account includes.

10. Or skip the browser setup

If your actual task is capturing a web page as an image or PDF, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a PNG, JPEG, WebP, or PDF. This is a screenshot workflow rather than a general HTML/XML-to-XLS or document-generation workflow.

Install nothing for this HTTP call; replace the placeholder with your API key. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await Bun.write('shot.webp', res);

For Node.js versions without Bun, write the response bytes with Node’s filesystem API:

import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000; every feature is on every plan.

Create a free ScreenshotNeo account to start with 1,000 screenshots per month and no card.

11. Frequently asked questions

Can I use DocRaptor for XLS or XLSX from Python?

Yes. The API reference lists PDF, XLS, and XLSX document types. Check the current format-specific requirements and consume the response using the file type your application expects.

Does test mode make production PDFs?

Test mode is for trial generation and its output is watermarked. Use the production setting for final documents and confirm account terms before deployment.

Should I choose DocRaptor or a screenshot API?

Choose DocRaptor when you need document conversion from HTML/XML or a URL, including spreadsheet formats. Choose a screenshot API when you need a rendered website image or capture-oriented PDF. ScreenshotNeo is one such option, with cleanup for consent banners and widgets and billing that excludes failed or blank captures.

Where do I find the exact async polling fields?

Use the current DocRaptor Python guide and API reference for the installed client version; the async response and callback flow should follow those documented fields.