How to Create a PDF Archive of Website Pages with PDFCrowd
Convert selected web pages to PDFs with PDFCrowd, then merge them into one archive. Includes cURL, Python, Node.js, layout options, and troubleshooting.
PDFCrowd can convert a reachable webpage URL to a PDF. To make an archive of several pages, collect the URLs, convert each one, and merge the resulting PDFs with PDFCrowd’s PDF-to-PDF API. The documented workflow is one URL per conversion request; the documentation reviewed here does not establish an automatic whole-site crawler or a one-request sitemap archive.
This guide uses PDFCrowd’s versioned HTTP API endpoint, https://api.pdfcrowd.com/convert/24.04/. Requests use HTTP Basic authentication and form fields. A successful conversion returns PDF bytes, so save the response as a binary file. See the [official HTTP API guide](https://pdfcrowd.com/api/html-to-pdf-http/) and [conversion examples](https://pdfcrowd.com/api/html-to-pdf-http/examples/).
1. Choose what belongs in the archive
First decide whether you need PDFs of public pages as PDFCrowd fetches them, pages represented by HTML already prepared in your application, or a page and its local assets. The input method matters:
| Input | Use it when | Important detail |
|---|---|---|
url |
The page is public and reachable from PDFCrowd’s servers. | It does not use your browser’s current login session or unsent form values. |
text |
Your application already has the HTML string to render. | Send the HTML as a form field. Include a <base href="..."> or absolute URLs for external assets if relative references need a base. |
file |
You have an HTML file or a packaged page and its resources. | For local assets, upload a supported archive containing the HTML and assets with their relative paths preserved. |
For an archive of multiple live pages, keep an explicit URL list and a stable ordering. A ZIP containing several HTML files is still an input for selecting an HTML document to convert; it does not automatically combine every page in that ZIP into one PDF. If the archive has multiple HTML files, set zip_main_filename to choose which one is converted. See the [HTTP API guide’s input and archive instructions](https://pdfcrowd.com/api/html-to-pdf-http/).
2. Convert each URL to a PDF
Use your PDFCrowd username as the HTTP Basic username and your API key as its password. The examples below read credentials from environment variables so they are not embedded in source code. Set PDFCROWD_USERNAME and PDFCROWD_API_KEY before running them. Do not commit credentials or print them in logs.
cURL: convert one page
curl --fail-with-body --silent --show-error \\
--user "$PDFCROWD_USERNAME:$PDFCROWD_API_KEY" \\
--output page-01.pdf \\
--form 'url=https://example.com/' \\
--form 'content_viewport_width=balanced' \\
'https://api.pdfcrowd.com/convert/24.04/'
content_viewport_width=balanced is an explicit example setting, not an API default. For a multi-page archive, repeat the request for each selected URL and save each response under a distinct filename. The [official cURL examples](https://pdfcrowd.com/api/html-to-pdf-http/examples/) show the request format and additional settings.
Python: convert a list of pages
Install the HTTP client with python -m pip install requests. Save this as archive_pages.py and run it after setting the two environment variables. It requests each URL separately, checks for HTTP errors, and writes response bytes without decoding them as text.
import os
from pathlib import Path
import requests
USERNAME = os.environ["PDFCROWD_USERNAME"]
API_KEY = os.environ["PDFCROWD_API_KEY"]
ENDPOINT = "https://api.pdfcrowd.com/convert/24.04/"
PAGES = [
("page-01.pdf", "https://example.com/"),
("page-02.pdf", "https://example.com/about"),
]
for filename, url in PAGES:
response = requests.post(
ENDPOINT,
auth=(USERNAME, API_KEY),
data={"url": url, "content_viewport_width": "balanced"},
timeout=(15, 180),
)
response.raise_for_status()
Path(filename).write_bytes(response.content)
print(f"Saved {filename} ({len(response.content)} bytes)")
The connect and read timeouts are client-side safeguards. Adjust them for your workload and network; a timeout does not establish whether a remote conversion completed, so handle retries carefully.
Node.js: convert a list of pages
This example uses Node.js with built-in fetch, FormData, and file APIs. It requires a Node.js release that provides those web APIs. Set the environment variables, save as archive-pages.mjs, then run node archive-pages.mjs.
import { writeFile } from "node:fs/promises";
const username = process.env.PDFCROWD_USERNAME;
const apiKey = process.env.PDFCROWD_API_KEY;
if (!username || !apiKey) {
throw new Error("Set PDFCROWD_USERNAME and PDFCROWD_API_KEY");
}
const endpoint = "https://api.pdfcrowd.com/convert/24.04/";
const pages = [
["page-01.pdf", "https://example.com/"],
["page-02.pdf", "https://example.com/about"],
];
const authorization = "Basic " + Buffer.from(`${username}:${apiKey}`).toString("base64");
for (const [filename, url] of pages) {
const form = new FormData();
form.set("url", url);
form.set("content_viewport_width", "balanced");
const response = await fetch(endpoint, {
method: "POST",
headers: { Authorization: authorization },
body: form,
signal: AbortSignal.timeout(180_000),
});
if (!response.ok) {
const detail = await response.text();
throw new Error(`PDFCrowd returned HTTP ${response.status}: ${detail}`);
}
const pdf = Buffer.from(await response.arrayBuffer());
await writeFile(filename, pdf);
console.log(`Saved ${filename} (${pdf.length} bytes)`);
}
Do not manually set the multipart Content-Type header in this example; the runtime adds the boundary used by FormData. The Authorization value is derived from the credential pair, so protect the environment and process logs accordingly.
3. Set layout and rendering options
PDF output is a print layout, not necessarily a screen-perfect snapshot. Choose settings deliberately and inspect the generated pages. Common HTTP form settings include:
| Goal | Settings to consider | What to check |
|---|---|---|
| Choose paper geometry | page_size, orientation, margin_top, margin_right, margin_bottom, margin_left; or no_margins |
Long lines, clipped content, and space for headers or footers. |
| Control responsive layout | content_viewport_width |
The browser viewport affects responsive breakpoints independently of the PDF paper size. |
| Add source and pagination | header_html, footer_html, header_height, footer_height |
Reserve adequate space and make sure the page content does not collide with them. |
| Apply print-specific styling | use_print_media=true, custom_css |
Print styles only help if the page provides them; custom CSS can hide navigation or adjust the content width. |
| Wait for dynamic content | wait_for_element or javascript_delay |
Prefer a selector that signals readiness when available; a fixed delay may be too short or waste time. |
| Include a particular region | element_to_convert |
Confirm the selector matches the content you intend to preserve. |
| Preserve document structure | Investigate the tagged-PDF setting in the API reference | Inspect the resulting file; do not assume a setting alone guarantees accessibility. |
Parameter names, accepted values, defaults, and constraints are maintained in the [HTML-to-PDF HTTP API reference](https://pdfcrowd.com/api/html-to-pdf-http/ref/). The API uses form fields rather than JSON request bodies; use multipart form data when uploading files. The endpoint is versioned, so keep the version explicit and check the current reference before changing it.
4. Merge the converted PDFs into one archive
After all conversions succeed, submit the files to the PDF-to-PDF API. Its documented join operation merges PDFs in the order they are added. The following cURL example combines three local files:
curl --fail-with-body --silent --show-error \\
--user "$PDFCROWD_USERNAME:$PDFCROWD_API_KEY" \\
--output website-archive.pdf \\
--form 'input_format=pdf' \\
--form 'f_1=@page-01.pdf' \\
--form 'f_2=@page-02.pdf' \\
--form 'f_3=@page-03.pdf' \\
'https://api.pdfcrowd.com/convert/24.04/'
Use an ordered manifest rather than relying on filesystem enumeration if page order matters. The [PDF-to-PDF API documentation](https://pdfcrowd.com/api/pdf-to-pdf-api/docs/) describes merging and other PDF operations; the [HTTP examples](https://pdfcrowd.com/api/pdf-to-pdf-http/examples/) show multipart inputs such as f_1 and f_2. If you need to preserve source URLs, add them to the converted pages using suitable header/footer settings or maintain a separate manifest.
5. Choose between URL mode and browser content
URL mode is the simplest option for public pages, but PDFCrowd fetches and renders the page on its servers. It does not inherit your local browser session. If a page requires authentication, configure source-site authentication separately, for example with website credentials, cookies, or a custom HTTP header. PDFCrowd API credentials authenticate your request to PDFCrowd; they are not credentials for the source website.
For content already loaded in a browser, PDFCrowd’s WebSave content mode can send the current HTML and relevant form values for conversion. This produces a new rendering on PDFCrowd’s servers, rather than a pixel capture of the browser screen. State that is not present in the submitted HTML—such as records never loaded by the page—will not appear automatically. See the [WebSave documentation](https://pdfcrowd.com/websave/).
Local pages such as http://localhost cannot be fetched from PDFCrowd’s servers. Upload the HTML and assets in a supported archive or send the HTML content instead, ensuring needed resources are included or reachable. For an archive with multiple HTML files, select the intended entry with zip_main_filename; that selects one document to convert and does not merge the rest.
6. Reliability, performance, and cost considerations
- Process pages independently. Keep one result file and status per URL. A failed page can then be retried without losing successfully converted pages.
- Use bounded concurrency. Parallel conversions can reduce total elapsed time, but increase simultaneous requests and resource use in your own application. Start conservatively and respect the account’s current service limits.
- Retry selectively. Retry transient network failures and temporary server errors with backoff. Do not blindly retry every authentication, invalid-option, or inaccessible-source error; correct the cause first.
- Validate each result. Check HTTP status before treating a response as a PDF. Keep conversion error bodies separate from output files, and consider checking the content type and whether the saved file begins with a PDF signature.
- Plan for changing pages. A URL conversion captures what the renderer receives at conversion time. If the source page changes, a later run may produce different content. Record the URL and capture time in your manifest.
- Estimate cost from current account terms. The dossier does not provide PDFCrowd pricing or a billing unit, so check the current plan and API terms before estimating a batch. A multi-page workflow makes multiple conversion requests plus a merge request.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 401 or 403 | PDFCrowd username or API key is wrong, missing, or malformed. | Check the credential pair and ensure the client sends HTTP Basic authentication. Keep PDFCrowd credentials distinct from source-site credentials. |
| Conversion fails for a URL | The URL is malformed, not public, blocked, or not reachable from PDFCrowd’s servers. | Use a complete http:// or https:// URL and verify server-side reachability. Configure separate source authentication if needed. |
| Localhost URL cannot be captured | localhost refers to the caller’s machine, not a host reachable by PDFCrowd. |
Send the HTML or upload it with its assets instead. |
| Output is an error message or a corrupt PDF | The client saved an error response as though it were PDF bytes. | Check status before saving or consuming output. Use --fail-with-body in cURL and inspect the error response. |
| Images or styles are missing | Relative asset URLs lack a base, local resources were not uploaded, or resources cannot be reached remotely. | Use absolute asset URLs or a <base> URL for HTML content; package local assets with their relative paths in an archive. |
| Dynamic content is absent | Rendering started before the page populated the content, or the content was never loaded into the submitted HTML. | Use wait_for_element for a reliable readiness marker or set an appropriate javascript_delay. For WebSave content mode, ensure the desired content is actually in the HTML sent. |
| Wrong layout or clipped content | Viewport, paper size, margins, orientation, or print CSS do not suit the page. | Adjust content_viewport_width, page dimensions, margins, print media, or custom CSS, then inspect the result. |
| Archive contains only one page | An HTML archive was mistaken for a multi-page PDF merge. | Convert each desired page to its own PDF, then use the PDF-to-PDF API with ordered f_1, f_2, and subsequent file fields. |
| Client reports timeout | The request took longer than the client timeout or the network failed while waiting. | Inspect the response if available, use a suitable timeout, and retry only when appropriate. Make retries safe for your workflow by tracking completed outputs. |
For diagnostics, PDFCrowd documents the errfmt=json query option for structured errors and debug_log=true for a conversion debug log. Capture the HTTP status, response headers, and error body when investigating. See the [official troubleshooting guidance](https://pdfcrowd.com/api/html-to-pdf-http/).
Or skip the browser setup
If you need a clean visual screenshot rather than a paginated PDF archive, [ScreenshotNeo](https://screenshotneo.com) provides a website screenshot API and MCP server. This one-call example returns a PDF; see the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).
curl -G "https://api.screenshotneo.com/v1/shot" \\
-d access_key=YOUR_API_KEY \\
--data-urlencode url=https://example.com/ \\
-d format=pdf \\
-o page.pdf
For ScreenshotNeo, cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. For a visual page-by-page archive, call it once per URL and organize the resulting files; the PDFCrowd conversion-and-merge workflow above is the documented route for combining selected website pages into one PDF.
Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card required.
FAQ
Can PDFCrowd create one PDF from a whole website automatically?
The workflow documented here converts chosen URLs individually and merges the PDFs. It does not establish automatic whole-site crawling.
Can I include the current values in a form?
URL mode does not share your browser’s form state. WebSave content mode can send current browser HTML and relevant form values, provided they are present in the submitted content.
Does merging PDFs preserve the order I selected?
The PDF-to-PDF API joins inputs in the order they are added. Assign file fields from an explicit ordered list.
Can I convert a private intranet page?
Only if PDFCrowd’s servers can reach it and the source authentication is configured appropriately. A page available solely through your local browser is not reachable by URL conversion.


