How to generate PDFs from a list of URLs with PDFCrowd
Convert a list of URLs to individual PDFs with PDFCrowd using cURL or Python. Learn how to name outputs, handle errors, set concurrency, and troubleshoot conversions.
To generate PDFs from a list of URLs with PDFCrowd, submit one conversion request per URL, save each successful response as its own PDF, and record the result against the original URL. PDFCrowd’s documented HTTP API accepts a single url form field per request; it does not convert an entire URL list in one request. The endpoint returns PDF bytes in the response. See the HTTP API reference.
The workflow is: prepare a URL list, give each URL a safe unique filename, make a conversion request for each item, save successful PDF bytes, and log failures so they can be retried independently. You can implement the loop with cURL, Python, the official SDK, or PDFCrowd’s command-line client.
What you need before converting
- A PDFCrowd account username and API key. These are used for HTTP Basic authentication.
- A plain-text or CSV file containing one source URL per line.
- A destination directory with enough space for the generated files.
- Network access from the machine running the script to
api.pdfcrowd.com.
The source page and its resources must also be reachable from PDFCrowd’s servers. A URL such as http://localhost:3000 points to the converter’s own machine, not yours. For private or local pages, use a reachable staging URL or submit HTML content using the API’s supported input options. Source-site login is separate from the credentials used to authenticate to PDFCrowd.
HTTP API: convert a URL list with cURL
The HTTP endpoint is https://api.pdfcrowd.com/convert/24.04/. Send a POST request with HTTP Basic authentication and form fields. Do not send JSON: the documented API expects form fields. A successful response body contains PDF bytes.
Convert one URL
curl --user 'YOUR_USERNAME:YOUR_API_KEY' \
--form 'url=https://example.com/' \
'https://api.pdfcrowd.com/convert/24.04/' \
--output 'example-com.pdf'
Replace the credentials and URL with your values. Keep API credentials out of source control and shell history where possible; supply them through a secret manager or environment-backed script in production.
Convert every URL in a text file
For a simple sequential batch, put one URL on each line of urls.txt. This shell script derives a readable filename from each host, adds an index to prevent collisions, saves each response, and appends a result to a tab-separated manifest. It continues after a failed request.
#!/usr/bin/env bash
set -u
: "${PDFCROWD_USERNAME:?Set PDFCROWD_USERNAME}"
: "${PDFCROWD_API_KEY:?Set PDFCROWD_API_KEY}"
endpoint='https://api.pdfcrowd.com/convert/24.04/'
mkdir -p pdfs
manifest='pdfs/manifest.tsv'
printf 'index\turl\tfile\tstatus\n' > "$manifest"
index=0
while IFS= read -r url || [[ -n "$url" ]]; do
[[ -z "$url" || "$url" == \#* ]] && continue
index=$((index + 1))
filename=$(printf 'page-%04d.pdf' "$index")
temp="pdfs/.${filename}.part"
error_file="pdfs/.${filename}.error"
if http_code=$(curl --silent --show-error \
--user "$PDFCROWD_USERNAME:$PDFCROWD_API_KEY" \
--form-string "url=$url" \
--output "$temp" \
--write-out '%{http_code}' \
"$endpoint" 2>"$error_file") && [[ "$http_code" == 200 ]]; then
mv "$temp" "pdfs/$filename"
printf '%s\t%s\t%s\tok\n' "$index" "$url" "$filename" >> "$manifest"
rm -f "$error_file"
else
rm -f "$temp"
printf '%s\t%s\t%s\tfail:%s\n' "$index" "$url" "$filename" "${http_code:-curl-error}" >> "$manifest"
printf 'Failed: %s (HTTP %s). See %s\n' "$url" "${http_code:-unknown}" "$error_file" >&2
fi
done < urls.txt
Set the credentials before running it, for example in the current shell session, then run the script. Review the manifest and error files after the batch. This example uses sequential requests, which is easy to reason about and avoids sending an unbounded burst. For robust production retries, use the retry policy described below rather than blindly repeating every failure.
Python: convert a list with PDFCrowd’s SDK
PDFCrowd documents a Python SDK. Install it with pip install pdfcrowd, then create an HtmlToPdfClient and call convertUrlToFile once per URL. The sample below writes each conversion to a unique numbered filename and keeps a CSV manifest with a status and error message.
import csv
import os
import time
from pathlib import Path
import pdfcrowd
USERNAME = os.environ["PDFCROWD_USERNAME"]
API_KEY = os.environ["PDFCROWD_API_KEY"]
INPUT = Path("urls.txt")
OUT = Path("pdfs")
OUT.mkdir(exist_ok=True)
client = pdfcrowd.HtmlToPdfClient(USERNAME, API_KEY)
with INPUT.open(encoding="utf-8") as source, (OUT / "manifest.csv").open(
"w", newline="", encoding="utf-8"
) as manifest_file:
writer = csv.DictWriter(
manifest_file, fieldnames=["index", "url", "file", "status", "error"]
)
writer.writeheader()
for index, raw in enumerate(source, start=1):
url = raw.strip()
if not url or url.startswith("#"):
continue
filename = f"page-{index:04d}.pdf"
destination = OUT / filename
temp = OUT / f".{filename}.part"
try:
# Each URL is a separate conversion request.
client.convertUrlToFile(url, str(temp))
temp.replace(destination)
writer.writerow({
"index": index, "url": url, "file": filename,
"status": "ok", "error": ""
})
except pdfcrowd.Error as exc:
temp.unlink(missing_ok=True)
writer.writerow({
"index": index, "url": url, "file": filename,
"status": "failed", "error": str(exc)
})
print(f"Failed: {url}: {exc}")
manifest_file.flush()
Set PDFCROWD_USERNAME and PDFCROWD_API_KEY in the environment before starting the script. The SDK’s exception type is pdfcrowd.Error; logging the exception per URL lets the batch continue without losing the status of earlier conversions. Consult the Python SDK guide for client methods and configuration.
Python with direct HTTP requests
If you prefer not to use the SDK, the API can also be called directly. This example uses requests, writes each response to a temporary file before renaming it, and records HTTP failures per URL.
import csv
import os
from pathlib import Path
import requests
ENDPOINT = "https://api.pdfcrowd.com/convert/24.04/"
AUTH = (os.environ["PDFCROWD_USERNAME"], os.environ["PDFCROWD_API_KEY"])
OUT = Path("pdfs")
OUT.mkdir(exist_ok=True)
with open("urls.txt", encoding="utf-8") as source, open(
OUT / "manifest.csv", "w", newline="", encoding="utf-8"
) as manifest_file:
writer = csv.DictWriter(
manifest_file, fieldnames=["index", "url", "file", "status", "detail"]
)
writer.writeheader()
for index, line in enumerate(source, start=1):
url = line.strip()
if not url or url.startswith("#"):
continue
filename = f"page-{index:04d}.pdf"
try:
response = requests.post(
ENDPOINT,
auth=AUTH,
data={"url": url},
timeout=75,
)
response.raise_for_status()
temp = OUT / f".{filename}.part"
temp.write_bytes(response.content)
temp.replace(OUT / filename)
writer.writerow({
"index": index, "url": url, "file": filename,
"status": "ok", "detail": ""
})
except requests.RequestException as exc:
writer.writerow({
"index": index, "url": url, "file": filename,
"status": "failed", "detail": str(exc)
})
print(f"Failed: {url}: {exc}")
manifest_file.flush()
The 75-second client timeout leaves room for response transfer but does not extend PDFCrowd’s server-side processing limit. To configure page rendering, use the documented form fields in the HTTP API or the corresponding SDK methods.
Node.js: submit one conversion per URL
The same request pattern works in Node.js using built-in fetch. This runnable script reads urls.txt, posts form-encoded fields with Basic authentication, saves the returned bytes, and writes a JSON Lines manifest. It processes requests sequentially for predictable load.
import { readFile, mkdir, writeFile, rename, appendFile, rm } from 'node:fs/promises';
const username = process.env.PDFCROWD_USERNAME;
const apiKey = process.env.PDFCROWD_API_KEY;
if (!username || !apiKey) throw new Error('Set PDFCROWD_USERNAME and PDFCROWD_API_KEY');
const endpoint = 'https://api.pdfcrowd.com/convert/24.04/';
const out = 'pdfs';
await mkdir(out, { recursive: true });
await writeFile(`${out}/manifest.jsonl`, '');
const urls = (await readFile('urls.txt', 'utf8'))
.split(/\r?\n/)
.map(line => line.trim())
.filter(line => line && !line.startsWith('#'));
for (let i = 0; i < urls.length; i++) {
const url = urls[i];
const filename = `page-${String(i + 1).padStart(4, '0')}.pdf`;
const temp = `${out}/.${filename}.part`;
const form = new URLSearchParams({ url });
const authorization = Buffer.from(`${username}:${apiKey}`).toString('base64');
try {
const response = await fetch(endpoint, {
method: 'POST',
headers: {
authorization: `Basic ${authorization}`,
'content-type': 'application/x-www-form-urlencoded',
},
body: form,
signal: AbortSignal.timeout(75000),
});
if (!response.ok) {
const detail = await response.text();
throw new Error(`HTTP ${response.status}: ${detail}`);
}
const bytes = Buffer.from(await response.arrayBuffer());
await writeFile(temp, bytes);
await rename(temp, `${out}/${filename}`);
await appendFile(`${out}/manifest.jsonl`, JSON.stringify({
url, file: filename, status: 'ok'
}) + '\n');
} catch (error) {
await rm(temp, { force: true });
await appendFile(`${out}/manifest.jsonl`, JSON.stringify({
url, file: filename, status: 'failed', error: String(error)
}) + '\n');
console.error(`Failed: ${url}: ${error}`);
}
}
Run with a Node.js version that provides global fetch and AbortSignal.timeout. Keep the API key in an environment variable or deployment secret.
Choosing filenames and preserving the URL mapping
Filename handling matters in batches because URLs can share a path, contain query parameters, or include characters that are unsafe on a filesystem. Prefer a stable index or generated identifier and keep the original URL in a manifest. If filenames must be human-readable, derive them from the host and a sanitized path, then append a short hash or index to avoid collisions.
- Do not use an untrusted URL directly as a filesystem path.
- Do not assume two different URLs map to different names after removing punctuation.
- Write to a temporary file and rename it only after a successful conversion, so a partial response is not mistaken for a finished PDF.
- Store status, HTTP or SDK error, source URL, output path, and attempt count in a CSV, JSON Lines file, or database.
Conversion options for page content and layout
The URL list controls which pages are converted. Rendering options control what each PDF looks like. Configure them consistently for the batch unless different URLs need different treatment. PDFCrowd documents options through its API and SDKs; use the exact parameter names and supported values from the HTTP API reference and developer API documentation.
| Need | Relevant setting or approach |
|---|---|
| Wait for dynamic content | Use wait_for_element or javascript_delay so client-rendered content has time to appear. |
| Use print-specific page styles | Enable use_print_media when the site has print CSS. |
| Correct an unexpectedly narrow or wide layout | Set content_viewport_width to the intended browser viewport width. |
| Access a protected source site | Provide source-site authentication, cookies, or headers separately from PDFCrowd Basic authentication. |
| Send your own HTML | Use the API’s HTML/content input options when the source cannot be fetched as a public URL. |
| Diagnose incomplete output | Inspect the response and optional debug log for resource download delays or JavaScript that exceeds the processing window. |
Do not assume that authentication to the conversion API also logs the converter into the source website. Configure source credentials as source-page inputs, and avoid placing secrets in URLs or logs.
Concurrency, retries, and batch reliability
Start sequentially unless you have a reason to parallelize. PDFCrowd says rate and concurrency limits depend on the license. Its documented status codes include 429 for request rate and 430 for requests already in progress. Check the limits that apply to your license before increasing workers; a large simultaneous burst can produce avoidable failures.
For a production batch, use a bounded worker pool and a bounded retry count. Retry temporary transport failures and transient service responses with increasing delays, for example 1, 2, and 4 seconds plus small random jitter. Do not retry authentication failures, invalid inputs, or repeatable layout problems without changing the cause. Log every attempt and retain the final failure in the manifest.
- Set a small concurrency limit appropriate to your license.
- Retry only likely transient errors, with a maximum attempt count.
- Use increasing delay between attempts; honor any retry guidance returned by the service.
- Write each finished output atomically and persist the manifest after each URL.
- On restart, skip records already marked successful and resume failed or unprocessed URLs.
PDFCrowd documents a maximum upload size of 300 MB and stops conversion work after 60 seconds of processing time. Those are vendor-stated operational limits, not a guarantee that every page will finish within that time. Increasing your client-side timeout cannot extend the server processing window. Slow resource downloads and long-running JavaScript can cause a conversion to hit the limit. See the API documentation.
Command-line client for scripted batches
If you already use shell scripts, scheduled jobs, or deployment automation, PDFCrowd also provides an HTML-to-PDF command-line client that accepts URL inputs. It can be a convenient interface for a loop while leaving filename mapping, logging, and retry behavior in your script. Check the command-line documentation for installation, authentication, and flags. The CLI and SDK are interfaces to the same hosted conversion service, so they do not represent different rendering engines.
Or skip the browser setup
If the deliverable can be a clean website screenshot rather than a paginated PDF, ScreenshotNeo can return PNG, JPEG, WebP, or PDF from one GET request. It is a website screenshot API and MCP server by Yorker Media. The same call works in cURL, Python, or Node.js; see the ScreenshotNeo API documentation for options and setup.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed, along with known newsletter popups and chat widgets, before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. The response headers report the page verdict and billing status.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 ScreenshotNeo shots a month, with no card required.
Troubleshooting PDFCrowd URL batches
| Symptom | Likely cause | What to do |
|---|---|---|
| Authentication error | Wrong PDFCrowd username or API key, or credentials not passed as Basic authentication. | Check account credentials and the request’s authentication configuration. Keep PDFCrowd credentials separate from source-site credentials. |
| API rejects the request body | JSON was sent instead of form fields, or the field name is incorrect. | Send a POST with form-encoded fields, including the singular url field. |
| Remote page cannot be fetched | The page is private, on localhost, behind a firewall, or otherwise unreachable from PDFCrowd’s servers. | Use a reachable URL or submit HTML content; configure source login separately if required. |
| HTTP 429 | The request rate exceeded the limit for the license. | Reduce request rate or concurrency, then retry with increasing delays. |
| HTTP 430 | Requests are already in progress beyond the allowed concurrency. | Lower the worker count and retry after current jobs finish. |
| Conversion runs too long or returns incomplete output | Slow resources or long-running JavaScript may exceed the 60-second server processing limit. | Inspect the response and debug log, simplify or speed up source dependencies, and use the documented wait controls only as needed. A longer client timeout does not change this limit. |
| Dynamic content is missing | The page renders content after the initial document load. | Configure wait_for_element or javascript_delay for the content that must appear. |
| PDF layout width is wrong | The converter viewport differs from the layout expected by the page. | Adjust content_viewport_width; check whether print styles should be enabled with use_print_media. |
| A saved file is not a valid PDF | An error response or interrupted response was saved as if it were successful. | Check the HTTP status before saving, inspect the error body, and write to a temporary file before renaming on success. |
| One failure stops the whole batch | The loop does not isolate errors per URL. | Catch errors inside the per-URL loop, append a failure record, and continue with the next item. |
Cost and runtime planning
The total work is one conversion request per URL, plus any retries. Runtime depends on page rendering, resource loading, your allowed request rate and concurrency, and output transfer. Avoid estimating a completion time from page count alone; measure a representative sample of your own URLs and plan for the slowest pages.
PDFCrowd’s pricing page describes a trial with 100 test API credits valid for one month and defines one credit as 0.5 MB of generated output. Credits, plan limits, and prices can change, so check the current plans page before estimating a production batch. The API reference also states the 300 MB maximum upload size and 60-second processing limit.
For cost control, avoid duplicate conversions when the source and settings have not changed, record completed outputs, and retry only errors likely to be temporary. If you need to rerun a batch, keep the URL list and conversion settings versioned so the result is reproducible.
Frequently asked questions
Can I send a list of URLs in one PDFCrowd request?
The reviewed HTTP API reference documents one url input per conversion request. Build the list workflow in your client by issuing a request for each URL.
Can PDFCrowd convert a page on my computer?
Not by fetching your machine’s localhost URL. The converter must be able to reach the source from its servers. Use a reachable page or provide the HTML using a supported content input.
Should I use the SDK, HTTP API, or CLI?
Use whichever fits your environment. The HTTP API gives direct control over requests, the Python SDK provides a language-specific client, and the CLI fits shell automation. All three are interfaces to the same service.
Does increasing my timeout fix a 60-second conversion limit?
No. A client timeout controls how long your program waits; it does not extend the documented server-side processing limit.


