How to Batch Convert Website URLs to PDF with PDFShift
Convert a list of website URLs to PDF with PDFShift using a reliable Python batch script, parallel webhooks, and practical guidance on limits and failures.
To batch convert URLs with PDFShift, send one POST request per URL to https://api.pdfshift.io/v3/convert/pdf, set the JSON source field to that URL, and authenticate with the X-API-Key header. Save each successful response as a PDF. For work that should run concurrently without holding open each request, PDFShift supports parallel, asynchronous conversions: provide a webhook and handle a completion callback for each source.
The examples below use Python for a reliable local batch, then show cURL and Node.js patterns. Keep your API key on a server or in a local environment variable; do not put it in browser-side code. See the PDFShift site and its webhook guide for current API details.
1. Prepare and validate the URL list
Store one absolute HTTP or HTTPS URL per line in urls.txt. Remove blank lines and duplicates, and ensure each address is one your conversion account is allowed to access. A URL that works in your logged-in browser may not work for a remote converter if it requires your browser session, local network access, or credentials.
https://example.com/
https://www.example.org/reports/annual
Keep the output directory separate from the input file. Use stable names based on the input index and host rather than the full URL: URLs can contain query strings, fragments, slashes, or sensitive tokens. The script below uses an index and sanitized hostname.
2. Batch convert with Python
Install the dependency with python -m pip install requests. Set the API key in the environment, then save this as batch_pdfshift.py. It processes URLs sequentially, writes each PDF to disk, records failures, and exits unsuccessfully if any item failed.
import os
import re
import sys
from pathlib import Path
from urllib.parse import urlparse
import requests
API_URL = "https://api.pdfshift.io/v3/convert/pdf"
API_KEY = os.environ.get("PDFSHIFT_API_KEY")
INPUT_FILE = Path("urls.txt")
OUTPUT_DIR = Path("pdfs")
TIMEOUT_SECONDS = 120
def safe_host(url: str) -> str:
host = urlparse(url).hostname or "website"
return re.sub(r"[^A-Za-z0-9.-]+", "_", host).strip("._") or "website"
def load_urls(path: Path) -> list[str]:
seen = set()
urls = []
for line in path.read_text(encoding="utf-8").splitlines():
url = line.strip()
if not url or url.startswith("#"):
continue
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.hostname:
print(f"Skipping invalid URL: {url}", file=sys.stderr)
continue
if url not in seen:
seen.add(url)
urls.append(url)
return urls
def main() -> int:
if not API_KEY:
print("Set PDFSHIFT_API_KEY before running.", file=sys.stderr)
return 2
if not INPUT_FILE.is_file():
print(f"Input file not found: {INPUT_FILE}", file=sys.stderr)
return 2
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
urls = load_urls(INPUT_FILE)
failures = []
with requests.Session() as session:
session.headers.update({"X-API-Key": API_KEY})
for index, url in enumerate(urls, start=1):
destination = OUTPUT_DIR / f"{index:04d}-{safe_host(url)}.pdf"
try:
response = session.post(
API_URL,
json={"source": url},
timeout=TIMEOUT_SECONDS,
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").lower()
if "application/pdf" not in content_type or not response.content.startswith(b"%PDF"):
raise ValueError(
f"Expected PDF bytes; received Content-Type {content_type!r}"
)
temporary = destination.with_suffix(".pdf.tmp")
temporary.write_bytes(response.content)
temporary.replace(destination)
print(f"Saved {url} -> {destination}")
except (requests.RequestException, OSError, ValueError) as exc:
failures.append((url, str(exc)))
print(f"Failed {url}: {exc}", file=sys.stderr)
if failures:
failure_file = OUTPUT_DIR / "failures.txt"
failure_file.write_text(
"\n".join(f"{url}\t{error}" for url, error in failures) + "\n",
encoding="utf-8",
)
print(f"{len(failures)} of {len(urls)} conversions failed; see {failure_file}.", file=sys.stderr)
return 1
print(f"Converted {len(urls)} URL(s).")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Run it with export PDFSHIFT_API_KEY='your-key' followed by python batch_pdfshift.py. On Windows PowerShell, use $env:PDFSHIFT_API_KEY='your-key'. The script writes each file through a temporary path and renames it after writing, reducing the chance that an interrupted run leaves a partial file with a final PDF name.
3. Use cURL for one URL or a small shell batch
A single conversion is a POST with a JSON body. The response is the PDF bytes when no filename-based hosted output is requested.
curl --fail-with-body --silent --show-error \
-X POST "https://api.pdfshift.io/v3/convert/pdf" \
-H "X-API-Key: $PDFSHIFT_API_KEY" \
-H "Content-Type: application/json" \
--data '{"source":"https://example.com/"}' \
-o example.pdf
For a simple shell list, loop over a file and stop or log errors deliberately. This example assigns an index to avoid unsafe filenames:
i=0
while IFS= read -r url || [ -n "$url" ]; do
[ -z "$url" ] && continue
i=$((i + 1))
printf 'Converting %s\n' "$url"
curl --fail-with-body --silent --show-error \
-X POST "https://api.pdfshift.io/v3/convert/pdf" \
-H "X-API-Key: $PDFSHIFT_API_KEY" \
-H "Content-Type: application/json" \
--data "$(python -c 'import json,sys; print(json.dumps({"source":sys.argv[1]}))' "$url")" \
-o "pdf-$i.pdf" || printf 'FAILED\t%s\n' "$url" >> failures.txt
done < urls.txt
For URLs with complex data, prefer a JSON library as in the Python example so quotes and special characters are escaped correctly. Do not enable shell tracing in a way that prints secrets into shared logs.
4. Convert URLs from Node.js
This runnable Node.js example uses the built-in fetch available in current Node versions. It reads urls.txt, validates basic URL syntax, checks the response type, and saves successful responses.
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import { URL } from 'node:url';
const key = process.env.PDFSHIFT_API_KEY;
if (!key) throw new Error('Set PDFSHIFT_API_KEY first');
await mkdir('pdfs', { recursive: true });
const lines = (await readFile('urls.txt', 'utf8')).split(/\r?\n/);
const urls = [...new Set(lines.map(x => x.trim()).filter(x => x && !x.startsWith('#')))];
const failures = [];
for (const [index, source] of urls.entries()) {
let parsed;
try { parsed = new URL(source); } catch { failures.push(`${source}\tInvalid URL`); continue; }
if (!['http:', 'https:'].includes(parsed.protocol)) {
failures.push(`${source}\tOnly HTTP and HTTPS URLs are supported`);
continue;
}
const host = (parsed.hostname.replace(/[^A-Za-z0-9.-]/g, '_') || 'website');
const path = `pdfs/${String(index + 1).padStart(4, '0')}-${host}.pdf`;
try {
const response = await fetch('https://api.pdfshift.io/v3/convert/pdf', {
method: 'POST',
headers: { 'X-API-Key': key, 'Content-Type': 'application/json' },
body: JSON.stringify({ source }),
signal: AbortSignal.timeout(120_000),
});
if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
const type = response.headers.get('content-type') || '';
const bytes = Buffer.from(await response.arrayBuffer());
if (!type.toLowerCase().includes('application/pdf') || !bytes.subarray(0, 4).equals(Buffer.from('%PDF'))) {
throw new Error(`Expected PDF; received ${type || 'unknown content type'}`);
}
await writeFile(path, bytes);
console.log(`Saved ${source} -> ${path}`);
} catch (error) {
failures.push(`${source}\t${error.message}`);
console.error(`Failed ${source}: ${error.message}`);
}
}
if (failures.length) {
await writeFile('pdfs/failures.txt', failures.join('\n') + '\n');
process.exitCode = 1;
}
5. Choose sequential requests or PDFShift parallel webhooks
Sequential processing is easiest to debug and naturally limits load, but takes longer for a large list. Parallel requests can finish sooner, but a large burst can overwhelm your own application or exceed the provider’s capacity. PDFShift documents a default parallel limit of 50 conversions across plans; contact PDFShift if you need more capacity. In its parallel mode, requests are accepted independently with HTTP 202, and each conversion result is sent to the configured webhook. [PDFShift webhook guide]
For webhook processing:
- Create a publicly reachable HTTPS endpoint that accepts PDFShift’s POST callback.
- Before submitting each job, save an internal job record containing your own job ID, the source URL, and intended output name.
- Include the webhook parameter in each conversion request. The callback should be matched to the correct input using the callback data and your stored job records; do not assume callbacks arrive in submission order.
- Validate callback requests according to the current PDFShift documentation, then persist the returned result or file to storage you control.
- Make callback handling idempotent: a repeated delivery must not create duplicate records or overwrite a newer result. Keep failures visible for manual or scheduled retry.
PDFShift’s guide describes the webhook parameter and completion POST. Confirm the current callback payload and authentication options in its API documentation before implementing a production receiver. [Webhook integration guide]
6. Options, output delivery, and document behavior
source: Set it to the URL to render. The API also supports HTML conversion; this guide focuses on website URLs.webhook: Use for asynchronous completion notifications when running parallel jobs.filename: In PDFShift’s Make guide, including a filename makes PDFShift store the generated file temporarily and return JSON with a URL instead of returning PDF bytes directly. That URL is available for two days; download it to durable storage if you need it longer. Omitting filename returns the PDF bytes. [PDFShift Make guide]- Other conversion settings: PDFShift documents options for page selection and other rendering behavior. Add only the parameters your conversion needs and check the current API reference for accepted names and values. [Page selection example]
If a page needs login, PDFShift’s guides show an auth option for basic authentication. Do not place credentials in URL query strings, and consider the sensitivity of sending page credentials to a third-party conversion service. [Secured page guide]
7. Limits, reliability, performance, and cost
PDFShift’s FAQ says its default parallel ceiling is 50 conversions, and its pricing page currently lists 50 free monthly credits, a 15 MB maximum file size, and a 30-second timeout on the free plan. The FAQ states default timeouts of 30 seconds for free and 100 seconds for paid plans; a timed-out conversion returns HTTP 408. Check the live plan details before planning a workload because quotas and limits can change. [PDFShift pricing]
Credits are measured by generated PDF size: PDFShift states one credit per 5 MB, with larger files consuming additional credits. The pricing page gives a 14 MB document as three credits. Estimate usage from output sizes, not just URL counts. [PDFShift pricing and credit accounting]
For dependable batches, cap concurrency below the service limit so you retain room for unrelated jobs and retries. Retry transient network errors and timeouts with a small bounded retry count and backoff; do not repeatedly retry a deterministic invalid URL or authentication error. Persist a manifest with each URL, status, output path, attempt count, and error. Checkpoint successes so restarting a batch does not reconvert every completed item. Use a longer client timeout than the provider’s conversion timeout, but recognize that increasing the client timeout cannot extend the provider’s configured limit.
PDFShift’s developer site says its conversion technology is HIPAA compliant and that documents are not stored unless specifically requested. These are vendor statements, not a determination that a specific workflow meets your obligations. For regulated or confidential content, verify the applicable agreement, configuration, data flow, and legal requirements before sending URLs or credentials. [PDFShift developer site]
8. Troubleshooting common batch failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Watermark on the PDF | The request was unauthenticated, the key was missing/incorrect, or sandbox mode was enabled. | Send the API key in X-API-Key and inspect the request headers. PDFShift says unauthenticated requests may fall back to watermarked output. [Authentication and watermark help] |
| 401 or 403 | Invalid key, wrong header, or key not loaded by the running process. | Check the environment variable and header spelling. PDFShift documents X-API-Key as its authentication header. |
| HTTP 202 but no PDF in the response | The job was accepted for asynchronous processing. | Use the webhook completion flow and track each job separately instead of trying to save the 202 body as a PDF. |
| HTTP 408 or client timeout | The page or its assets took too long, or the service timeout was reached. | Check whether the URL is responsive without a session, reduce unnecessary page complexity, and review the timeout for your plan. Retrying once with backoff may help transient slowness; repeated retries do not fix a consistently slow page. |
| Response saved but file is not a PDF | An error response or JSON callback body was written as though it were PDF data. | Check HTTP status, content type, and the %PDF file signature before saving. |
| Some callbacks do not match a URL | The implementation assumed callbacks arrive in order or did not persist job metadata. | Store a job record at submission time and correlate each callback using its documented fields; treat callbacks as independent events. |
| Hosted PDF link has expired | The Make guide’s temporary file URL was not downloaded within its two-day availability period. | Download the result promptly and save it in storage you control. |
| PDF is incomplete or missing images | The website serves assets slowly, blocks remote requests, or renders differently to a remote browser. | Check the source page’s accessibility and rendering behavior, then consult PDFShift’s current rendering options. Validate representative output before processing a large batch. |
Or skip the browser setup
If you need a screenshot-style capture as well as PDFs, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return PNG, JPEG, WebP, or PDF. For a PDF capture, add the documented PDF format parameter:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o shot.pdf
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account to get started.
FAQ
Can I batch convert hundreds of URLs at once?
Yes, by submitting one conversion per URL. For asynchronous parallel processing, PDFShift documents a default ceiling of 50 simultaneous conversions; larger capacity requires contacting the vendor.
Does every URL cost one PDFShift credit?
Not necessarily. Credits are charged by generated size: one credit covers up to 5 MB, and larger output uses more.
Do I need to keep the generated PDFs on PDFShift?
No. The documented direct-response pattern returns PDF bytes for your application to save. The hosted URL workflow described in the Make guide is temporary, so download files needed for longer retention.
Can I use the same PDFShift batch script for private pages?
Only if the conversion service can access the page using a supported authentication method. Confirm the current secured-page options and assess whether sending those credentials to an external service is appropriate.


