How to Create an HTML Page with PhantomJS
Create an HTML page from a string or URL with PhantomJS, inspect it, render an image, troubleshoot failures, and understand its legacy limits.
Direct answer: use page.setContent(html, url) when your HTML is a string, and page.open(url, callback) when the page already exists at a URL. Use page.evaluate() to inspect the document and page.render() to save a visual result. PhantomJS must be started from the command line with a JavaScript file, and the script must call phantom.exit() when it finishes.
PhantomJS is archived and its 2.x branch is deprecated. The instructions below are useful for maintaining an existing script or reproducing a legacy workflow; they are not a recommendation to start a new production project. The project wiki documents the deprecated status, and the GitHub repository was archived on May 30, 2023.
Choose the input method
| Your input | Method | What it does |
|---|---|---|
| An HTML string | page.setContent(html, url) |
Creates the page context from supplied markup, sets its URL, and reloads it without making an HTTP request. |
| A web address | page.open(url, callback) |
Loads the URL and reports success or fail in the callback. |
| Document data | page.evaluate(fn) |
Runs JavaScript inside the page and returns simple, serializable values. |
| A visual file | page.render(path) |
Saves the rendered page as an image. |
Prerequisites and the command-line workflow
- Install a PhantomJS 2.x binary that is compatible with the legacy system you are maintaining.
- Save JavaScript to a file such as
create-page.js. - Run it with
phantomjs create-page.js. - Wait for asynchronous page work to finish, then call
phantom.exit().
The official quick start introduces the webpage module and the command-line execution model: PhantomJS Quick Start and Web Page Module API.
Create a page from an HTML string
This is the literal answer when you already have HTML text. The second argument gives the page a base URL, which matters for relative links, stylesheets, images, and scripts.
var page = require('webpage').create();
var html = '<!doctype html>' +
'<html>' +
'<head>' +
' <meta charset="utf-8">' +
' <title>Example page</title>' +
' <style>body { font-family: sans-serif; margin: 40px; }</style>' +
'</head>' +
'<body>' +
' <h1>Created with PhantomJS</h1>' +
' <p id="status">Ready</p>' +
'</body>' +
'</html>';
page.setContent(html, 'http://example.com/');
var details = page.evaluate(function () {
return {
title: document.title,
heading: document.querySelector('h1').textContent,
status: document.getElementById('status').textContent
};
});
console.log(JSON.stringify(details));
page.render('example.png');
phantom.exit();
setContent() sets both the content and URL and reloads the page without an HTTP request. See the setContent API reference.
Why the base URL matters
Markup such as <img src="images/logo.png"> or <script src="app.js"> is resolved relative to the URL passed to setContent(). Use a URL that represents the origin you expect. If all resources are embedded directly in the HTML, the URL still provides a useful document location and origin.
Load an existing URL
Use page.open() when PhantomJS should perform the navigation. Always check the callback status before reading the DOM or rendering.
var page = require('webpage').create();
var url = 'https://example.com/';
page.open(url, function (status) {
if (status !== 'success') {
console.error('Could not load ' + url + ' (status: ' + status + ')');
phantom.exit(1);
return;
}
var result = page.evaluate(function () {
return {
title: document.title,
links: document.querySelectorAll('a').length
};
});
console.log(JSON.stringify(result));
page.render('example-url.png');
phantom.exit();
});
The callback receives success or fail; the documented method is described in the page.open API reference.
Inspect and modify the document with evaluate()
page.evaluate() runs in the webpage context, so it can read the DOM and change it before a render. Only simple primitives and JSON-serializable objects cross the boundary. Functions, closures, and DOM nodes cannot be returned directly.
var page = require('webpage').create();
page.setContent(
'<html><body><h1 id="title">Old title</h1></body></html>',
'http://example.com/'
);
page.evaluate(function () {
document.getElementById('title').textContent = 'New title';
document.body.style.backgroundColor = '#f4f4f4';
});
var title = page.evaluate(function () {
return document.getElementById('title').textContent;
});
console.log(title);
page.render('modified.png');
phantom.exit();
See the official evaluate() documentation for the execution boundary and return-value rules.
Wait for asynchronous content before rendering
Pages that use timers or client-side rendering may not be ready when open() returns. Poll for a condition with setInterval, or use a bounded delay. Always include a timeout so a broken page cannot keep the process alive forever.
var page = require('webpage').create();
var waited = 0;
var maxWait = 10000;
page.open('https://example.com/app', function (status) {
if (status !== 'success') {
console.error('Navigation failed');
phantom.exit(1);
return;
}
var timer = setInterval(function () {
waited += 250;
var ready = page.evaluate(function () {
return document.querySelector('#app-ready') !== null;
});
if (ready || waited >= maxWait) {
clearInterval(timer);
if (!ready) {
console.error('Timed out waiting for #app-ready');
phantom.exit(1);
return;
}
page.render('app.png');
phantom.exit();
}
}, 250);
});
Rendering and output details
- Render after readiness: call
page.render('file.png')only after navigation and required JavaScript have completed. - Choose a viewport: set
page.viewportSize = { width: 1280, height: 800 };before loading if responsive layout matters. - Capture the whole document: inspect
page.scrollSizeand assign it topage.viewportSizebefore rendering when a full-page image is required. - Keep output deterministic: use fixed viewport dimensions, a fixed base URL, predictable data, and an explicit wait condition.
var page = require('webpage').create();
page.viewportSize = { width: 1280, height: 800 };
page.open('https://example.com/', function (status) {
if (status !== 'success') {
phantom.exit(1);
return;
}
var size = page.evaluate(function () {
return {
width: Math.max(document.body.scrollWidth, document.documentElement.scrollWidth),
height: Math.max(document.body.scrollHeight, document.documentElement.scrollHeight)
};
});
page.viewportSize = { width: size.width, height: size.height };
page.render('full-page.png');
phantom.exit();
});
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
page is undefined |
The webpage module was not created. | Start with var page = require('webpage').create();. |
| Blank or partial image | Rendering happened before asynchronous content finished. | Poll for a DOM marker or wait with a timeout before calling render(). |
status === 'fail' |
Navigation, DNS, TLS, or the remote server failed. | Log the URL and status, verify it outside PhantomJS, and stop before evaluating the page. |
| Relative assets do not load | The HTML string has no useful base URL. | Pass an appropriate second argument to setContent(). |
| Evaluation returns an empty or invalid value | A DOM node, function, or non-serializable object crossed the boundary. | Return strings, numbers, booleans, arrays, or plain JSON objects. |
| Process never exits | phantom.exit() was omitted or a timer remains active. |
Clear timers and call phantom.exit() on success and failure paths. |
| Modern site behaves incorrectly | PhantomJS uses an old browser engine and is no longer maintained. | Treat the script as legacy maintenance work and migrate to a maintained browser automation stack or a screenshot API. |
Performance, reliability, and cost notes
- Performance: avoid unnecessary waits, load only the assets required for the capture, and reuse a predictable HTML template. Full-page rendering costs more time and memory than a fixed viewport.
- Reliability: check every navigation result, wait on a page-specific readiness signal, bound all waits, and record the URL and failure status for debugging.
- Compatibility: modern JavaScript, TLS behavior, browser APIs, and sites that actively block old automation may not work reliably in PhantomJS.
- Cost: the PhantomJS workflow has no API request fee, but you maintain the binary, runtime, browser compatibility, retries, and infrastructure yourself. A hosted service shifts those operational tasks to an API call.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Before capture, it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await Bun.write('shot.webp', image);
Create a free ScreenshotNeo account to get 1,000 screenshots each month with no card.
FAQ
Can PhantomJS create a page without requesting a URL?
Yes. Pass an HTML string to page.setContent(). The method reloads the page context without making an HTTP request.
When should I use open() instead?
Use page.open() when PhantomJS should navigate to an existing URL and load its resources.
Can evaluate() return an element?
No. Return serializable data such as an element’s text, attributes, or a plain object containing those values.
Why does PhantomJS need phantom.exit()?
The command-line process remains alive while timers or page activity exist. Call phantom.exit() after success and failure handling.
Is PhantomJS suitable for a new application?
The project is archived and the 2.x branch is deprecated. Use these instructions for legacy compatibility, then plan migration to maintained tooling.


