ScreenshotNeo

BlogHTML to image & PDF

How to use PDFCrowd to convert a web page to PDF in Python

Convert a web page URL to PDF with PDFCrowd’s Python client, choose file or in-memory output, tune page layout, and diagnose common errors.

By the ScreenshotNeo team4 October 20266 min read

Use PDFCrowd’s official Python client: install the pdfcrowd package, create an HtmlToPdfClient with your username and API key, then call convertUrlToFile() to save the result. The same client can return PDF bytes with convertUrl() or write to a stream with convertUrlToStream(). The URL must use HTTP or HTTPS. See PDFCrowd’s Python guide and API reference.

1. Install the PDFCrowd Python client

python -m pip install pdfcrowd

Run the command in the environment that will execute your script. If you use a virtual environment, activate it first. The official package wraps PDFCrowd’s HTML-to-PDF API.

2. Convert a URL directly to a PDF file

Save this as convert_page.py. Replace the credential placeholders with your account values before running it.

import os
import sys
import pdfcrowd

URL = "https://www.example.com"
OUTPUT_PATH = "example.pdf"

username = os.environ.get("PDFCROWD_USERNAME")
api_key = os.environ.get("PDFCROWD_API_KEY")
if not username or not api_key:
    raise SystemExit("Set PDFCROWD_USERNAME and PDFCROWD_API_KEY first")

try:
    client = pdfcrowd.HtmlToPdfClient(username, api_key)
    client.convertUrlToFile(URL, OUTPUT_PATH)
    print(f"Saved PDF to {OUTPUT_PATH}")
except pdfcrowd.Error as error:
    sys.stderr.write(f"PDFCrowd conversion failed: {error}\n")
    raise

Set credentials in your shell rather than committing them to source control:

export PDFCROWD_USERNAME="YOUR_USERNAME"
export PDFCROWD_API_KEY="YOUR_API_KEY"
python convert_page.py

PDFCrowd’s guide uses demo/demo for trying its examples. Those are demo credentials; use your own account credentials for production. The guide describes a free trial or API license for production use. See the official setup guide.

3. Choose how your application receives the PDF

Choose the output method based on what happens after conversion. These methods all convert a URL; they differ in where the resulting PDF goes.

Method Use it when Result
convertUrlToFile(url, path) A script should save a named file PDF written directly to the given path
convertUrl(url) You need to return a PDF from a web endpoint or process it in memory PDF bytes
convertUrlToStream(url, out_stream) You want to write to a supplied stream PDF written to that stream

Return PDF bytes

import os
import pdfcrowd

client = pdfcrowd.HtmlToPdfClient(
    os.environ["PDFCROWD_USERNAME"],
    os.environ["PDFCROWD_API_KEY"],
)
pdf_bytes = client.convertUrl("https://www.example.com")

with open("example.pdf", "wb") as output:
    output.write(pdf_bytes)

Because this method holds the response in memory, use it when the PDF must be passed to another component, such as an HTTP response, or further processed. For large outputs or direct file saving, consider the file or stream method.

Write to a stream

import os
import pdfcrowd

client = pdfcrowd.HtmlToPdfClient(
    os.environ["PDFCROWD_USERNAME"],
    os.environ["PDFCROWD_API_KEY"],
)
with open("example.pdf", "wb") as output:
    client.convertUrlToStream("https://www.example.com", output)

PDFCrowd documents these methods in its Python reference.

4. Adjust responsive layout and print settings

A page’s PDF layout depends on the viewport used to render the page and the PDF page settings. PDFCrowd’s examples document content viewport width, page size, orientation, and margins. Start with viewport width when a responsive page looks too narrow, too wide, or different from its browser layout; then adjust the paper settings to suit the intended print result. There is no single setting that fits every site.

import os
import pdfcrowd

client = pdfcrowd.HtmlToPdfClient(
    os.environ["PDFCROWD_USERNAME"],
    os.environ["PDFCROWD_API_KEY"],
)
client.setContentViewportWidth("960px")
client.setPageSize("A4")
client.setOrientation("landscape")
client.setMarginTop("10mm")
client.setMarginRight("10mm")
client.setMarginBottom("10mm")
client.setMarginLeft("10mm")
client.convertUrlToFile("https://www.example.com", "example.pdf")

For a responsive page, try client.setContentViewportWidth("balanced") as shown in PDFCrowd’s examples, or use a specific width such as 960px. The example settings illustrate available controls; they do not guarantee a particular target site’s layout. Check the resulting PDF and adjust based on the content.

5. Troubleshoot and inspect failures

Catch pdfcrowd.Error around the conversion so the failure is visible to your script. PDFCrowd’s examples also show enabling debug logging and inspecting conversion metadata.

import os
import sys
import pdfcrowd

client = pdfcrowd.HtmlToPdfClient(
    os.environ["PDFCROWD_USERNAME"],
    os.environ["PDFCROWD_API_KEY"],
)
client.setDebugLog(True)

try:
    client.convertUrlToFile("https://www.example.com", "example.pdf")
except pdfcrowd.Error as error:
    sys.stderr.write(f"Conversion error: {error}\n")
    try:
        print("Debug log URL:", client.getDebugLogUrl())
        print("Job ID:", client.getJobId())
        print("Pages:", client.getPageCount())
        print("Output bytes:", client.getOutputSize())
        print("Credits consumed:", client.getConsumedCredits())
        print("Credits remaining:", client.getRemainingCredits())
    except pdfcrowd.Error as diagnostic_error:
        sys.stderr.write(f"Could not read diagnostics: {diagnostic_error}\n")
    raise
Symptom Likely cause What to check
Authentication error Credentials are missing, mistyped, or not the intended account credentials Confirm the username and API key values passed to HtmlToPdfClient; do not assume demo credentials work for production.
URL rejected or page not retrieved The input is not a publicly reachable HTTP or HTTPS URL, or the scheme is unsupported Use an absolute http:// or https:// URL and verify that the target is reachable by the conversion service.
PDF layout is clipped or too narrow The rendered content viewport or page settings do not match the site’s responsive layout Adjust content viewport width, page size, orientation, and margins; inspect the output after each change.
No output file appears The conversion raised an exception, or the process cannot write to the chosen path Print or log the caught pdfcrowd.Error, confirm the output directory is writable, and use an absolute path to remove ambiguity.
Script reports an import error The package was installed into a different Python environment Install with python -m pip install pdfcrowd using the same python executable that runs the script.

6. Reliability, performance, and cost considerations

  • Keep credentials private. Load them from environment variables or a secret store, especially in deployed applications.
  • Handle exceptions at the job boundary. Log the error and PDFCrowd diagnostic identifiers where available so a failed conversion can be investigated.
  • Choose output mode deliberately. Returning bytes uses application memory for the result; direct file and stream methods fit file-oriented workflows.
  • Account for remote conversion time. URL conversion depends on a remote API and on retrieving the target page. Avoid blocking a latency-sensitive request for long conversions; use your application’s background-job pattern when appropriate.
  • Check current plan and credit details. The cited setup and examples establish the API workflow and diagnostics, but do not provide enough information to state current prices or conversion limits. Review PDFCrowd’s current account terms before estimating production cost.

7. cURL, Python, and Node.js alternatives

The task here is specifically PDFCrowd’s Python client. Its official recipe is the Python code above. The research sources for this article document that client and its API, but do not provide a verified cURL or Node.js PDFCrowd example, so a runnable version for those clients should be taken from PDFCrowd’s current HTTP API documentation rather than guessed. The same documentation describes an HTTP API and its parameters: PDFCrowd HTML-to-PDF HTTP API.

Or skip the browser setup

If the goal is a visual capture rather than a paginated PDF, ScreenshotNeo returns a screenshot from one GET request. For a full PDF workflow, PDFCrowd’s conversion methods above are the relevant tool; ScreenshotNeo is the alternative to try first when a clean page image is what you need.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

See the ScreenshotNeo API docs for parameters. It removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots. One thousand screenshots a month are free with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Can PDFCrowd convert a local HTML file with this URL method?

The documented URL conversion methods accept HTTP or HTTPS URLs. This recipe is for a web page address; use the appropriate documented file or HTML-input method if your source is local content.

Which conversion method should a Flask or Django view use?

Use convertUrl() when your view needs PDF bytes to return in its response. Use a file or stream workflow when the application should persist the result instead.

Does setting a viewport width guarantee the PDF matches a browser screenshot?

No. Viewport width and page settings are controls for layout, but the target page’s responsive behavior determines how it renders. Inspect the output and tune the documented settings for that page.

References