ScreenshotNeo

BlogHTML to image & PDF

How to Generate a PDF with DocRaptor from a Private Web Page

Generate a PDF from login-protected content by retrieving it in your application server and sending its HTML to DocRaptor. Learn the authentication limits and safer alternatives.

By the ScreenshotNeo team4 October 20269 min read

For a page behind your application’s login, the clearest general workflow is to authenticate and authorize the user in your own application, retrieve or construct the page’s HTML on your server, then send that HTML to DocRaptor as document_content in a server-side request. Do not assume DocRaptor receives the user’s browser cookies or login session.

DocRaptor accepts either HTML content or a URL as the document source. Its documentation lists HTTP Basic Auth options and discusses authentication for protected assets, but the references reviewed do not clearly establish whether those options authenticate the initial document_url fetch. Confirm that behavior with DocRaptor before depending on it for private pages. DocRaptor API documentation

1. Choose how to provide the private page

Approach When it fits Key consideration
Server retrieves HTML, then submits document_content Your application can access the content after authenticating and authorizing the user. Most defensible general pattern for private content. Keep the DocRaptor key on your server.
Submit document_url The converter can retrieve the page and its required assets. A URL that works in a user’s browser may not be reachable to DocRaptor. Confirm the primary URL authentication method with the vendor.
Configured referrer-based endpoint You control the publishing domain and can configure it in DocRaptor. This documented flow checks an authorized HTTP referrer; it is not documented as forwarding a visitor’s login session.

In each case, authorize the requesting user in your application before creating a document. A user who can request a PDF should not gain access to another user’s page by changing an ID or URL.

2. Server-side implementation with document_content

Retrieve the protected content using your application’s existing authorization, then post the resulting HTML to DocRaptor. The example below uses Python and the requests package. Set secrets in the process environment or a secret manager; do not put a live API key in browser JavaScript.

import os
import requests
from flask import Flask, abort, request, send_file
from io import BytesIO

app = Flask(__name__)

DOCRAPTOR_API_KEY = os.environ["DOCRAPTOR_API_KEY"]
DOCRAPTOR_URL = "https://docraptor.com/docs"

@app.post("/reports/<report_id>.pdf")
def report_pdf(report_id):
    # Replace these placeholders with your application's authentication,
    # authorization, and database lookup.
    user = require_authenticated_user(request)
    report = load_report(report_id)
    if report is None or not user_can_view(user, report):
        abort(404)

    # Build HTML from authorized application data. Escape any untrusted
    # values with your template engine rather than concatenating raw input.
    html = render_report_html(report)

    payload = {
        "user_credentials": {
            "username": DOCRAPTOR_API_KEY,
            "password": ""
        },
        "doc": {
            "document_content": html,
            "name": f"report-{report_id}.pdf",
            "type": "pdf",
            "test": True
        }
    }
    response = requests.post(DOCRAPTOR_URL, json=payload, timeout=90)
    response.raise_for_status()

    return send_file(
        BytesIO(response.content),
        mimetype="application/pdf",
        as_attachment=True,
        download_name=f"report-{report_id}.pdf"
    )

Confirm the current request schema and authentication format against the DocRaptor API reference before deploying. The example uses test mode (test: true), which produces a watermarked PDF; change it only after validating layout and resource access with non-sensitive content. Replace the placeholder application functions with your real session, permission, lookup, and template code.

For production, also set sensible request timeouts, limit concurrent conversions, validate report identifiers, and return a controlled error to the user if conversion fails. Avoid logging API credentials, private HTML, or PDF bytes.

3. cURL, Python, and Node.js request patterns

These examples show the DocRaptor document request shape for HTML already obtained by a trusted server. They use a placeholder HTML document and test mode. Check the current API reference for request schema changes before integrating.

cURL

curl --user "$DOCRAPTOR_API_KEY:" \
  --header "Content-Type: application/json" \
  --data '{"doc":{"document_content":"<html><body><h1>Private report</h1></body></html>","name":"report.pdf","type":"pdf","test":true}}' \
  https://docraptor.com/docs \
  --output report.pdf

Python

import os
import requests

api_key = os.environ["DOCRAPTOR_API_KEY"]
payload = {
    "doc": {
        "document_content": "<html><body><h1>Private report</h1></body></html>",
        "name": "report.pdf",
        "type": "pdf",
        "test": True,
    }
}
response = requests.post(
    "https://docraptor.com/docs",
    auth=(api_key, ""),
    json=payload,
    timeout=90,
)
response.raise_for_status()
with open("report.pdf", "wb") as output:
    output.write(response.content)

Node.js

import { writeFile } from "node:fs/promises";

const apiKey = process.env.DOCRAPTOR_API_KEY;
if (!apiKey) throw new Error("Set DOCRAPTOR_API_KEY");

const payload = {
  doc: {
    document_content: "<html><body><h1>Private report</h1></body></html>",
    name: "report.pdf",
    type: "pdf",
    test: true
  }
};
const basic = Buffer.from(`${apiKey}:`).toString("base64");
const response = await fetch("https://docraptor.com/docs", {
  method: "POST",
  headers: {
    "Authorization": `Basic ${basic}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify(payload),
  signal: AbortSignal.timeout(90000)
});
if (!response.ok) {
  throw new Error(`DocRaptor request failed: ${response.status} ${await response.text()}`);
}
await writeFile("report.pdf", Buffer.from(await response.arrayBuffer()));

DocRaptor documents API-key authentication and document_content as an input. Keep requests on a server you control. Do not ship the key in a frontend bundle or call the API directly from public browser code.

4. What to know about document_url authentication

Use document_url only if DocRaptor can retrieve the document and every required asset. Your server can make a private URL reachable through a controlled mechanism, but that mechanism must be designed explicitly. For example, an application can create a short-lived, narrowly scoped URL if its security model permits it; protect it against reuse and avoid embedding sensitive data in query strings that could enter logs.

The API reference lists prince_options[http_user] and prince_options[http_password]. The security white paper discusses HTTP Basic Auth for protected assets. The reviewed references do not unambiguously say those options authenticate the initial document URL. Treat that as an open question and ask DocRaptor or validate it using a controlled, non-sensitive endpoint before relying on it.

Do not forward a user’s session cookie to a conversion service without explicitly reviewing the security implications and the vendor’s supported behavior. A browser session cookie is often broader and longer-lived than the narrowly scoped access needed for a single PDF.

5. The referrer-based flow for a site you control

DocRaptor documents a from_site endpoint for configured domains: configure the domain in the account and link to https://api.docraptor.com/docs/from_site from a page whose HTTP referrer contains the authorized domain. This avoids placing an API key in that link flow. Follow DocRaptor’s current instructions for the required link and settings.

This is a site-owner integration, not a general way to make DocRaptor inherit a visitor’s private application session. Review how the protected page authorizes the request, and do not assume the referrer alone grants access to private content. See DocRaptor’s documentation.

6. JavaScript, styles, and linked assets

  • JavaScript: DocRaptor says JavaScript is disabled by default. Enable javascript when the page depends on client-side rendering, and check whether the separate Prince JavaScript option is needed. Dynamic applications may need a deliberate completion signal or wait behavior. Test the actual page because enabling script execution alone does not guarantee asynchronous content has finished.
  • CSS, fonts, and images: These may be separate network requests. Verify their URLs and access independently; successful HTML submission does not prove the assets loaded.
  • Protected assets: The security white paper describes HTTPS, HTTP Basic Auth for assets, limited or short-lived asset URLs, IP allowlisting, and a customer-controlled secured proxy as controls to consider. Select a method appropriate to your threat model and confirm current vendor support.
  • Resource errors: The API reference includes ignore_resource_errors. While diagnosing missing content, avoid suppressing errors until you understand which resource failed.

References: API documentation, JavaScript guide, and DocRaptor’s security information.

7. Troubleshooting

Symptom Likely cause What to check or change
The PDF shows a login page The converter fetched the page without the application’s browser session, or received login HTML as the document. Inspect the HTML sent to DocRaptor. Authenticate and authorize in your application, then submit the authorized HTML with document_content.
DocRaptor cannot retrieve the URL The URL is private, network-restricted, or the assumed authentication method does not apply to the initial fetch. Verify reachability and confirm primary URL authentication behavior with DocRaptor. Consider server-side retrieval and content submission.
Styles, images, or fonts are missing Assets use relative, inaccessible, authenticated, or invalid URLs. Inspect each asset URL and its access controls independently. Check HTTPS, use appropriately scoped asset access, and examine resource errors.
Charts or dynamic sections are blank JavaScript is disabled or asynchronous rendering had not completed. Enable the required JavaScript option and test the page’s rendering completion behavior.
Some resource failures are not visible Resource errors may be ignored. Review ignore_resource_errors and enable useful diagnostics while troubleshooting.
API key appears in page source The request is being made from browser-side JavaScript. Move the request to your server and read the key from server-side configuration. DocRaptor warns that its browser JavaScript library exposes the key.
PDF layout differs from the page Print styles, unavailable fonts, dynamic content, or unsupported page assumptions affect rendering. Test with representative non-sensitive content, check linked resources, and tune the document’s print CSS and rendering options.

8. Security, performance, and cost

Security

  • Perform authorization for every PDF request, including object-level access checks.
  • Keep the API key in server-side configuration, rotate it according to your policy, and restrict access to the service that needs it.
  • Send only the data needed to render the PDF. Review current DocRaptor security terms and your own retention and data-handling requirements.
  • Protect the output as carefully as the source page. Generated PDFs may contain the same private information even after the HTML request is complete.
  • Use test mode and non-sensitive content while validating. DocRaptor documents test mode as producing watermarked PDFs; check the API reference for current limits.

Performance and reliability

  • Rendering time depends on the document, scripts, and reachable assets; the cited workflow documentation provides no relevant benchmark to promise a conversion time.
  • Set a client timeout, handle non-success responses, and give users a retry path that does not accidentally create duplicate work. Use bounded concurrency and queue longer jobs if your application needs to protect request workers.
  • Keep the HTML and asset set deterministic where possible. A page depending on live data or third-party resources can render differently or fail when those resources change or become unavailable.
  • Record a request identifier and status for support without logging the API key, private source HTML, or sensitive PDF output.

Cost

Check DocRaptor’s current pricing and plan terms directly before estimating operating cost. This article makes no pricing or conversion-cost claim. Include retries, test conversions, and any storage or delivery costs in your application’s own estimate.

9. Or skip the browser setup

If your goal is to capture a webpage as an image or PDF rather than build and maintain a browser-rendering pipeline, ScreenshotNeo is a website screenshot API and MCP server. Its one-request API can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for options and integration details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. An MCP server lets AI agents use screenshot tools. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. For a private page, first confirm that the URL can be accessed through a supported, secure mechanism; never place credentials in a public URL.

Sign up free for 1,000 screenshots a month, no card required.

10. Frequently asked questions

The reviewed documentation does not say that it forwards a visitor’s browser session. Arrange authorized content delivery explicitly from your application.

Can I use DocRaptor without exposing my API key?

Yes. Make the API request from a trusted server, or evaluate the documented referrer-based flow for a domain you control. Do not put a live key in browser code.

Is document_content safe for sensitive pages?

It is an API input option, not a blanket security guarantee. Send only authorized data and review the vendor’s current security terms and your organization’s requirements.

Should I turn on JavaScript for every PDF?

No. It is disabled by default; enable it when the document needs client-side rendering, then verify dynamic content with representative pages.

Sources