How to Print a Full-Page Website to PDF with PDFCrowd
Use PDFCrowd to save a whole website as a PDF. Choose a single tall page or a conventional paginated document, then tune its layout.
Short answer: PDFCrowd converts a webpage URL to PDF through its HTML-to-PDF HTTP API. To make one vertically expanded page that contains the whole page, set page_height=-1. To make a conventional PDF that can be printed on standard paper, choose a paper size such as A4 or Letter and let the content flow across multiple pages. These are different meanings of “full-page”: one tall page versus all page content across normal sheets.
The examples below use PDFCrowd’s documented HTTP API. Its versioned endpoint accepts form fields, uses HTTP Basic authentication with your PDFCrowd username and API key, and returns PDF bytes on success. See the HTTP API guide and parameter reference.
1. Choose one tall PDF page or standard pages
| Output | Set | Useful when | Trade-off |
|---|---|---|---|
| One expanded page | page_height=-1, with an intentional page width |
Reading a long infographic or preserving one continuous vertical layout | The page can be extremely tall and awkward to print or open in some PDF viewers. |
| Conventional paginated PDF | A standard page_size, such as A4 or Letter |
Printing, sharing, or reading a document in familiar page-sized sections | Content flows across pages and may split elements at page boundaries. |
PDFCrowd documents a safe maximum of 200 inches for custom page heights and warns that larger dimensions may fail to open in some viewers. The special -1 value expands the page vertically to fit the content. If printability matters, start with standard pages instead.
2. Create a PDF from a URL with cURL
Replace the demo credentials with your PDFCrowd username and API key. This example requests a standard A4 PDF, with a balanced content viewport and no page margins:
curl -f -s -S \
-u 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
-o page.pdf \
-F 'url=https://example.com/' \
-F 'page_size=A4' \
-F 'content_viewport_width=balanced' \
-F 'no_margins=true' \
https://api.pdfcrowd.com/convert/24.04/
For one long page, replace the page_size field with page_height=-1. Keep an explicit page width if you need to control how wide the tall page is:
curl -f -s -S \
-u 'PDFCROWD_USERNAME:PDFCROWD_API_KEY' \
-o page.pdf \
-F 'url=https://example.com/' \
-F 'page_width=8.27in' \
-F 'page_height=-1' \
-F 'content_viewport_width=balanced' \
-F 'no_margins=true' \
https://api.pdfcrowd.com/convert/24.04/
The endpoint and request format are documented by PDFCrowd. Form fields must be sent as form data, not JSON. The successful response is binary PDF content, so save it directly to a file instead of trying to parse it as text or JSON.
3. Match the webpage layout to the PDF
The browser’s content viewport and the PDF paper size are separate controls. The viewport width affects responsive breakpoints: a narrow viewport may activate a mobile layout, while a wide one may produce desktop navigation and columns. The resulting PDF page size then determines how that rendered layout fits on paper.
- Page size: The documented default is A4. The reference lists A0 through A6 and Letter. A4 is a common international choice; Letter is common in the United States.
- Orientation: Use portrait for ordinary reading and landscape where the content is naturally wide. Orientation matters for standard page sizes.
- Margins: Set
margin_top,margin_right,margin_bottom, andmargin_leftfor explicit page margins, or useno_marginsfor edge-to-edge output. Headers and footers also need room if enabled. - Print styling: If the site has print-specific CSS, enable
use_print_media. Otherwise,custom_csscan hide navigation, ads, or sidebars or adjust spacing. - Page range: Use
print_page_rangewhen only selected pages of a paginated result are needed.
Whitespace can come from two layers. PDF page margins are set by the conversion options; CSS margins and padding on the source document are part of the webpage itself. If blank borders remain after reducing PDF margins, inspect the page’s html and body styling. PDFCrowd’s page-layout FAQ describes this distinction.
4. Use Python to save the PDF bytes
This runnable example uses the Python standard library. Set credentials through environment variables so they are not embedded in source code. It checks the HTTP status and content type before writing the response:
import os
import urllib.parse
import urllib.request
username = os.environ["PDFCROWD_USERNAME"]
api_key = os.environ["PDFCROWD_API_KEY"]
endpoint = "https://api.pdfcrowd.com/convert/24.04/"
fields = {
"url": "https://example.com/",
"page_size": "A4",
"content_viewport_width": "balanced",
"no_margins": "true",
}
body = urllib.parse.urlencode(fields).encode("utf-8")
request = urllib.request.Request(endpoint, data=body, method="POST")
request.add_header("Content-Type", "application/x-www-form-urlencoded")
password_mgr = urllib.request.HTTPPasswordMgrWithDefaultRealm()
password_mgr.add_password(None, endpoint, username, api_key)
opener = urllib.request.build_opener(urllib.request.HTTPBasicAuthHandler(password_mgr))
try:
with opener.open(request, timeout=90) as response:
content_type = response.headers.get("Content-Type", "")
pdf_bytes = response.read()
if not content_type.startswith("application/pdf"):
raise RuntimeError(f"Expected a PDF response, got {content_type!r}")
with open("page.pdf", "wb") as output:
output.write(pdf_bytes)
except urllib.error.HTTPError as error:
details = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"PDFCrowd returned HTTP {error.code}: {details}") from error
print("Saved page.pdf")
For a tall single page, change the fields dictionary to include "page_width": "8.27in" and "page_height": "-1", and remove page_size. Use the same form-field names documented in the parameter reference.
5. Use Node.js to save the PDF bytes
This example uses built-in Node.js APIs and sends form-encoded data with Basic authentication:
const username = process.env.PDFCROWD_USERNAME;
const apiKey = process.env.PDFCROWD_API_KEY;
if (!username || !apiKey) throw new Error('Set PDFCROWD_USERNAME and PDFCROWD_API_KEY');
const endpoint = 'https://api.pdfcrowd.com/convert/24.04/';
const form = new URLSearchParams({
url: 'https://example.com/',
page_size: 'A4',
content_viewport_width: 'balanced',
no_margins: 'true',
});
const credentials = Buffer.from(`${username}:${apiKey}`).toString('base64');
const response = await fetch(endpoint, {
method: 'POST',
headers: {
Authorization: `Basic ${credentials}`,
'Content-Type': 'application/x-www-form-urlencoded',
},
body: form,
signal: AbortSignal.timeout(90000),
});
if (!response.ok) {
throw new Error(`PDFCrowd returned HTTP ${response.status}: ${await response.text()}`);
}
const contentType = response.headers.get('content-type') || '';
if (!contentType.startsWith('application/pdf')) {
throw new Error(`Expected a PDF response, got ${contentType}`);
}
const { writeFile } = await import('node:fs/promises');
await writeFile('page.pdf', Buffer.from(await response.arrayBuffer()));
console.log('Saved page.pdf');
For a tall page, send page_width and page_height=-1 instead of page_size. Do not expose the API key in browser-side JavaScript: keep authenticated conversion requests on a server you control.
6. Tune waiting and page cleanup
A URL conversion loads the page and its resources from PDFCrowd’s servers. If the page populates content after its initial load, use wait_for_element for a known selector or javascript_delay for a measured delay. Neither setting guarantees that every interactive state or late-loading element will be ready; inspect the output and the debug log.
Use custom_css to hide page furniture such as a navigation bar or cookie overlay when appropriate. Use custom_javascript to adjust content before conversion if needed. Selectors and scripts must match the actual source page. Avoid long waits and scripts that continue running: they add latency and can cause a conversion to exceed the service processing limit.
7. Other input types and edge cases
- HTML file: Upload the file in the
filefield. A local path typed as a string is not uploaded content. - HTML string: Send rendered markup in the
textfield. Use absolute URLs for remote CSS and images, or add a suitable<base>element. - Local assets: Package the HTML and related assets in a supported archive while preserving relative paths. If the archive has multiple HTML files, choose the entry with
zip_main_filename. - Local development server: PDFCrowd cannot fetch your computer’s
localhost. Upload the HTML or send its contents instead. - Protected source page: PDFCrowd’s Basic credentials authenticate you to the conversion API. They do not log you into the source website. Configure source-site credentials, cookies, or headers separately where applicable.
- Content behind scripts or iframes: Confirm the page’s state can be rendered and check the relevant wait and iframe settings. Some content may depend on interaction or permissions unavailable during conversion.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 400 response or “no input” | Missing input field or JSON body instead of form data | Send exactly one of url, text, or file as form data; inspect the error reason code. |
| 401 response | Missing or incorrect API credentials, or inactive license | Check the PDFCrowd username and API key. These are different from source-site credentials. |
| 403 response | Account suspended or credits unavailable | Check account state and remaining credits. |
| 413 response | Uploaded input exceeds the documented 300 MB limit | Reduce or repackage the input. |
| 429 or 430 response | Request-rate limit or concurrent-request limit reached | Slow submissions for 429; allow active requests to finish for 430. |
| 503 or transient network failure | Temporary service or network problem | Retry a bounded number of times with increasing delays. |
| Conversion processing timeout | Slow assets or long-running JavaScript; the documented processing limit is 60 seconds | Reduce page work, fix slow resource URLs, or reduce wait duration. A larger client timeout cannot fix a server-side processing timeout. |
| Images or CSS missing | Resources are inaccessible, relative paths do not resolve, or they require authentication | Use accessible absolute URLs, set a base URL, authenticate resource requests where applicable, or upload an archive with local assets. |
| Login screen appears instead of content | The source website requires its own login | Supply source-site authentication or cookies separately. |
| Dynamic content is absent | Capture occurred before the content appeared, or the page requires interaction | Wait for a selector or use a short delay, then review debug output. Do not assume a delay recreates user interaction. |
| Layout is unexpectedly narrow or wide | Content viewport selected a different responsive breakpoint | Adjust content_viewport_width independently of the PDF paper size. |
| Unexpected whitespace | Source CSS margins/padding remain after PDF margins are removed | Inspect both PDF margin settings and the source document’s CSS. |
| Output saved but is not a usable PDF | An error body was saved as if it were PDF bytes | Check HTTP status and content type before writing; use -f or --fail-with-body in cURL. |
For diagnosis, preserve the HTTP status, response headers, and body. Add debug_log=true to a request and use ?errfmt=json on the endpoint for structured errors. PDFCrowd’s HTTP guide documents useful response headers, including job ID, reason code, consumed credits, page count, and output size.
9. Performance, reliability, and cost considerations
Conversion time depends on page rendering, JavaScript, and resource downloads. A practical request should use an HTTP timeout long enough for a conversion and file transfer, while remembering the service’s documented 60-second processing limit. Avoid launching an unbounded number of conversions: PDFCrowd says rate and concurrency limits depend on the license, and exposes 429 and 430 responses for those limits.
For resilient automation, distinguish temporary failures from permanent ones. Retry transient network errors or 503 responses a bounded number of times with increasing delays. Do not automatically retry invalid requests, authentication failures, account/credit errors, or oversized uploads without first correcting the cause. Record the job ID and reason code when present so repeated failures can be traced.
PDFCrowd responses include credit-related headers when present; record consumed and remaining credits to monitor usage. The dossier does not establish a universal per-conversion cost, so check the current plan and account terms before estimating a production workload. Reduce unnecessary page resources and excessive waits to limit both conversion time and avoidable use.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Use its PDF option when you want a PDF without managing a browser conversion setup. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The one-call example above returns an image by default; use ScreenshotNeo’s PDF output option when the required result is a PDF. Cookie banners are accepted like a visitor and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers say which occurred. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month, with no card required.
FAQ
Does page_height=-1 create multiple pages?
It requests one page that expands vertically to fit the document. Use a standard paper size if you need ordinary pagination.
Can PDFCrowd fetch a URL that only exists on my laptop?
No. Its servers cannot reach your local localhost. Send HTML content or upload the document and its assets instead.
Why is the PDF still padded with blank space when I set zero margins?
Zero PDF page margins do not remove whitespace added by the webpage’s own CSS. Check the source document’s html and body margins and padding too.
Can I print only part of a page?
For selected pages of a paginated PDF, consult print_page_range in the parameter reference. To convert a particular page element, see element_to_convert.


