How to Render and Download PDFs with PhantomJS
Render a webpage to PDF with PhantomJS, control paper size and margins, handle supplied HTML, and distinguish rendering from downloading an existing PDF.
Short answer: PhantomJS renders the currently loaded webpage with page.render('output.pdf'). Load a URL with page.open(), wait for a successful callback, set page.paperSize when you need controlled dimensions, and then render. This creates a new PDF from the page. It does not download an existing PDF response unchanged.
PhantomJS 2.1 is the latest stable release. Development is suspended, and the project repository has been archived read-only since May 30, 2023. Treat the examples below as legacy-maintenance guidance and test every target page before relying on the output in a new workflow. See the official repository.
Render a webpage to PDF
The basic sequence is:
- Create a WebPage object.
- Optionally set
paperSize. - Call
page.open(url, callback). - Render only when the callback status is
success. - Exit PhantomJS after the file is written.
var page = require('webpage').create();
page.paperSize = {
format: 'A4',
orientation: 'portrait',
margin: '1cm'
};
page.open('https://example.com', function (status) {
if (status !== 'success') {
console.log('Unable to load page');
phantom.exit(1);
return;
}
page.render('output.pdf');
phantom.exit();
});
Save this as render.js and run it with the PhantomJS executable:
phantomjs render.js
The documented page.render method selects the output format from the filename extension. A filename ending in .pdf produces PDF output. The example above is a minimal API illustration; it was not tested against every website or PhantomJS build.
Control paper size, orientation and margins
Set page.paperSize before calling render when the document needs a predictable layout. The paperSize API supports named formats and explicit dimensions.
| Setting | Values and behavior |
|---|---|
format |
A3, A4, A5, Legal, Letter or Tabloid. |
orientation |
portrait or landscape; portrait is the default. |
width, height |
Explicit dimensions using mm, cm, in or px. A missing unit is interpreted as pixels. |
margin |
One value for all sides, or an object with top, left, bottom and right. The default margin is zero. |
header, footer |
Optional repeating header or footer definitions with a height and callback-generated contents. |
page.paperSize = {
width: '210mm',
height: '297mm',
margin: {
top: '15mm',
right: '12mm',
bottom: '15mm',
left: '12mm'
}
};
For PDF output, the quality argument is not a PDF-quality control. It applies to JPEG and PNG output. Page dimensions and margins are the controls that affect PDF layout.
Render HTML that is already in memory
If the HTML is supplied by your program, use page.setContent(html, baseUrl) instead of making PhantomJS navigate to a URL. The method loads the supplied content without making an HTTP request and sets the current URL. A meaningful base URL helps relative stylesheets, images and links resolve in practice.
var page = require('webpage').create();
var html = '' +
'<html><head><style>body { font-family: sans-serif; }</style></head>' +
'<body><h1>Invoice</h1><p>Generated from a string.</p></body></html>';
page.paperSize = {
format: 'A4',
orientation: 'portrait',
margin: '1cm'
};
page.setContent(html, 'https://example.com/');
page.render('invoice.pdf');
phantom.exit();
See the setContent documentation for the method signature.
Wait for pages that render asynchronously
page.open reports that navigation completed, but a modern application may still fetch data or update the DOM afterward. The sources document the open callback and rendering API; they do not guarantee that every site’s client-side work has finished when the callback fires.
For a known element, poll for a readiness marker before rendering:
var page = require('webpage').create();
page.paperSize = { format: 'A4', margin: '1cm' };
page.open('https://example.com/report', function (status) {
if (status !== 'success') {
phantom.exit(1);
return;
}
var attempts = 0;
var timer = setInterval(function () {
var ready = page.evaluate(function () {
return !!document.querySelector('[data-report-ready="true"]');
});
if (ready) {
clearInterval(timer);
page.render('report.pdf');
phantom.exit();
} else if (++attempts >= 20) {
clearInterval(timer);
console.log('Timed out waiting for report data');
phantom.exit(2);
}
}, 250);
});
Use a marker your application controls when possible. A fixed delay can be a fallback, but it may be too short on a slow run and unnecessarily long on a fast one.
Render versus download: two different jobs
These terms describe different operations:
| Goal | Correct approach |
|---|---|
| Create a PDF from a webpage | Use PhantomJS page.open, then page.render('file.pdf'). |
| Save a PDF that a server already returns | Use an HTTP client or file-download code and handle redirects, authentication, headers and the response body. |
page.render writes a rendering of the current page. It is not documented as a way to retrieve an existing PDF response byte-for-byte.
Download an existing PDF with cURL
curl -L --fail --output report.pdf \
'https://example.com/report.pdf'
For an authenticated endpoint, add the authentication mechanism required by that service, such as a header or cookie, and verify the response status and content type.
Download an existing PDF with Python
import requests
url = "https://example.com/report.pdf"
with requests.get(url, stream=True, timeout=90) as response:
response.raise_for_status()
with open("report.pdf", "wb") as output:
for chunk in response.iter_content(chunk_size=1024 * 1024):
if chunk:
output.write(chunk)
Download an existing PDF with Node.js
const fs = require('node:fs');
const response = await fetch('https://example.com/report.pdf', {
redirect: 'follow'
});
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
const file = fs.createWriteStream('report.pdf');
file.write(Buffer.from(await response.arrayBuffer()));
file.end();
Backgrounds, assets and legacy-browser limits
The official FAQ warns that a page without a defined background may render transparently. Set a background on the page or its body when the PDF must have a solid background:
page.evaluate(function () {
document.body.style.backgroundColor = '#ffffff';
});
Images, stylesheets and fonts must be reachable by the PhantomJS process. Relative URLs depend on the document URL; HTML supplied through setContent should therefore use an appropriate base URL or absolute asset URLs. Cross-origin requests, unsupported JavaScript features and layout differences in the old WebKit engine can all change the result. Inspect the generated PDF rather than assuming a current browser and PhantomJS will match.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
Unable to load page |
DNS, TLS, redirect, network or server failure. | Log the callback status, verify the URL from the same machine, and handle redirects and authentication. |
| Blank or incomplete PDF | Rendering started before client-side content was ready. | Wait for a DOM readiness marker or a carefully chosen delay before calling render. |
| Transparent background | The page does not define a background color. | Set body or page background CSS before rendering. |
| Wrong page size or clipping | paperSize was omitted or dimensions and margins conflict with the layout. |
Set the format or explicit dimensions before render; review margins and orientation. |
| Missing images or CSS | Relative URLs, inaccessible assets or failed requests. | Provide a useful base URL, use absolute paths where appropriate, and check asset availability. |
| Modern site behaves differently | PhantomJS uses an old WebKit engine and is no longer developed. | Test the exact pages, simplify unsupported client-side code where possible, or evaluate a maintained browser automation tool for new reliability-sensitive work. |
| Expected source PDF was changed | page.render rendered a webpage instead of downloading the PDF response. |
Use an HTTP/file download path for an existing PDF URL. |
Performance, reliability and cost considerations
- Performance: Rendering time depends on navigation, scripts, assets and any readiness wait. Avoid arbitrary long delays; wait for a specific condition when you can.
- Reliability: Check the
success/failstatus, set an outer process timeout, and preserve non-zero exit codes for failed jobs. Keep representative pages in a regression set because the archived engine can diverge from current browser behavior. - Output validation: Confirm that the file exists and has the expected size. For downloads, validate HTTP status and response headers separately from PDF parsing.
- Cost: PhantomJS itself does not provide a hosted rendering service or usage meter. Your costs are the machines, storage and operational work needed to run it. A hosted API can remove browser setup but introduces its own plan and request limits.
Or skip the browser setup
ScreenshotNeo is a website screenshot API that can return PNG, JPEG, WebP or PDF from one GET request. It removes cookie banners, newsletter popups and chat widgets before capture; bot checks, blank pages and failed loads are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for the complete option list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is available on every plan. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Official PhantomJS references
page.openfor URL loading and callback status.page.renderfor writing rendered output.page.paperSizefor dimensions, orientation, margins, headers and footers.page.setContentfor in-memory HTML.- The official
rasterize.jsexample for an end-to-end rendering example.
FAQ
Can PhantomJS convert an existing PDF URL?
Use an HTTP client to download an existing PDF. PhantomJS rendering is for producing a PDF from the current webpage.
Does page.render need a PDF-specific quality setting?
No. For PDF output, configure paper size, orientation and margins. The quality argument concerns JPEG and PNG output.
What happens if page.open fails?
Do not render. Log the failure, return a non-zero exit code, and investigate networking, redirects, TLS, authentication or server behavior.
Is PhantomJS suitable for a new production renderer?
It can maintain an existing workflow, but development is suspended and the repository is archived. Test carefully and assess a maintained browser automation option for new systems.


