ScreenshotNeo

BlogHow-to

How to Scrape Dynamic Page Content With PhantomJS

Learn how to extract JavaScript-rendered content with PhantomJS, handle evaluate serialization, diagnose missing data, and choose a maintained alternative.

By the ScreenshotNeo team30 September 20268 min read

How to Scrape Dynamic Page Content With PhantomJS

PhantomJS can scrape content that exists only after JavaScript runs. The basic pattern is: create a webpage object, call page.open, verify the callback status, and use page.evaluate to inspect the rendered DOM. Return only strings, numbers, booleans, arrays, or plain objects that can cross PhantomJS’s JSON serialization boundary.

This approach is mainly useful when you are maintaining an existing PhantomJS job. PhantomJS development is suspended, its GitHub repository has been archived, and the project wiki describes the 2.x line as deprecated and no longer maintained. For a new scraper, plan a migration to a maintained browser automation tool. The extraction concepts in this guide still help you understand why a legacy script misses dynamic content.

1. The minimal PhantomJS scraper

Save this as scrape.js and run it with the PhantomJS executable:

var webpage = require('webpage');
var page = webpage.create();
var url = 'https://example.com';

page.open(url, function (status) {
  if (status !== 'success') {
    console.log('Could not load page: ' + status);
    phantom.exit(1);
    return;
  }

  var result = page.evaluate(function () {
    var heading = document.querySelector('h1');
    return {
      title: document.title,
      heading: heading ? heading.innerText : '',
      url: location.href
    };
  });

  console.log(JSON.stringify(result));
  phantom.exit();
});

page.open opens the URL and invokes its callback with a page status such as success or fail. The official reference documents this behavior at the page.open API. A successful callback means the page load completed from PhantomJS’s perspective; it does not prove that a single-page application has finished every later asynchronous request.

page.evaluate executes a function inside the loaded page context, as described in the official evaluate API. That is where document.querySelector, textContent, attributes, and other DOM operations belong.

2. Why dynamic content is missing

Traditional HTML scraping reads the response body. A JavaScript application can return a nearly empty document, then fetch data and build the visible elements later. PhantomJS provides a browser context, so the DOM inspected by evaluate can include those rendered elements. The difficult part is deciding when the application state you need is ready.

Dynamic pages may populate the DOM after the initial load callback.
Dynamic pages may populate the DOM after the initial load callback.
  • Load completion is not application readiness. The page.open callback reports the load status documented by PhantomJS, not a universal “all JavaScript is finished” signal.
  • Choose a meaningful readiness condition. Identify a selector or state that only appears when the target data is present, such as .product-card, [data-loaded="true"], or a non-empty results container.
  • Do not treat an arbitrary sleep as universal. A fixed delay can be too short on a slow run and wasteful on a fast one. If you retain PhantomJS, implement a site-specific readiness check using the APIs available in your existing version and verify it against the target site.

The example below keeps extraction focused on a serializable result. Replace the selectors with selectors from the page you own or are authorized to crawl.

var webpage = require('webpage');
var page = webpage.create();

page.open('https://example.com/catalog', function (status) {
  if (status !== 'success') {
    console.log(JSON.stringify({ error: 'page_open_failed', status: status }));
    phantom.exit(1);
    return;
  }

  // Confirm the target site's ready state before calling evaluate.
  // The exact waiting strategy is site-specific.
  var products = page.evaluate(function () {
    var nodes = document.querySelectorAll('.product-card');
    var rows = [];
    for (var i = 0; i < nodes.length; i++) {
      var node = nodes[i];
      var name = node.querySelector('.product-name');
      var price = node.querySelector('.price');
      rows.push({
        name: name ? name.textContent.trim() : '',
        price: price ? price.textContent.trim() : '',
        href: node.querySelector('a') ? node.querySelector('a').href : ''
      });
    }
    return rows;
  });

  console.log(JSON.stringify(products));
  phantom.exit();
});

3. The evaluate serialization boundary

Arguments passed into evaluate and values returned from it cross between the PhantomJS script and the web page through JSON-compatible serialization. Return primitive values or plain data structures. Do not return a DOM node, a function, a window object, or an object containing circular references.

Safe to return Problematic to return Use instead
String, number, boolean document.querySelector('h1') Return node.innerText or an attribute
Array of plain objects A closure or function Compute the value inside evaluate
Null or simple object Window, document, or circular object Copy only the fields required by the scraper

Keep selectors and extraction logic inside the evaluated function. Keep file writes, process control, and final JSON output outside it. For example:

var data = page.evaluate(function () {
  var meta = document.querySelector('meta[name="description"]');
  return {
    title: document.title,
    description: meta ? meta.getAttribute('content') : null
  };
});

console.log(JSON.stringify(data));

A console.log executed inside the page is page-context output. It will not necessarily appear in the PhantomJS process output unless you configure page.onConsoleMessage. Returning structured data is more predictable for a scraper.

4. A practical extraction checklist

  1. Confirm that the target permits your access and that your request rate is appropriate.
  2. Open the URL and handle a fail status immediately.
  3. Inspect the page in a normal browser and identify the selector that represents completed data.
  4. Wait for that site-specific state before extracting.
  5. Use page.evaluate to select only the fields you need.
  6. Convert every field to a serializable value, handling missing elements with empty strings or null.
  7. Serialize outside the page context and exit with a non-zero status on operational failure.
  8. Log enough context to distinguish navigation failure, an empty result, and a selector change.

5. Common errors and fixes

“Could not load page” or a fail status

Cause: DNS, TLS, connectivity, navigation, or a page-level failure prevented PhantomJS from completing the open operation.

Fix: Log the status, verify the URL from the same machine, and retry according to your job’s policy. Do not call evaluate as if a usable document were available.

The result is empty, but the browser shows content

Cause: Extraction ran before the application inserted its data, or the selector changed.

Fix: Inspect the rendered DOM, choose a readiness selector tied to the content, and update selectors. A successful page.open callback alone is insufficient for every asynchronous application.

evaluate returns null or unusable data

Cause: The function returned a DOM node, a function, or another value that cannot cross the serialization boundary.

Fix: Return text, numbers, booleans, arrays, or a plain object. Extract innerText, textContent, href, and other scalar fields inside the evaluated function.

Page-console diagnostics never appear

Cause: Logs from the page context are separate from the outer PhantomJS process.

Fix: Return diagnostic values from evaluate, or configure the page’s console-message handler in the legacy script.

The script works locally but fails in production

Cause: Different PhantomJS versions, network policies, certificates, user-agent behavior, or timing expose assumptions in the script.

Fix: Record the runtime version and URL, make readiness checks explicit, bound retries, and keep selectors under review. PhantomJS’s suspended maintenance increases compatibility risk as sites adopt newer browser features.

6. Performance, reliability, and cost considerations

Dynamic rendering is more expensive than downloading static HTML because a browser process must parse, execute, and lay out the page. Reduce work by extracting only required fields, avoiding unnecessary navigation, and reusing a page where your legacy job safely permits it. Keep concurrency conservative: too many simultaneous PhantomJS processes can exhaust memory and make timing less predictable.

Reliability depends on the target application’s state, not only on the network. Use a readiness condition tied to the data, distinguish an empty legitimate result from a failed load, and capture structured error records. Keep a migration plan because the PhantomJS repository is archived and the project README states that development is suspended. The project wiki also describes the 2.x branch as deprecated and no longer maintained: repository and wiki.

Running PhantomJS yourself has infrastructure costs: process startup, memory, patching, browser binaries, and operational retries. A hosted screenshot or rendering API can move that browser work behind an HTTP request, but compare billing rules and failure handling carefully.

7. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. You can request a PNG, JPEG, WebP, or PDF with one GET request, while still controlling the capture details that commonly require browser automation.

A capture service can clean common overlays before billing for the result.
A capture service can clean common overlays before billing for the result.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo documentation for request parameters and response details. Its capture options include full-page screenshots with lazy images loaded, element capture by CSS selector, dark mode, 12 device presets or any viewport, retina scale, custom CSS and JavaScript, click-before-capture actions, selector or network-idle waits, hidden selectors, blocked ads and trackers, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, image resizing, configurable caching TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. PDF requests support paper size, margins, landscape mode, and page ranges.

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Each response reports the result through X-Page-Verdict and X-Billed headers. An MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

There is a free allowance of 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account to try the API.

8. When to keep PhantomJS and when to migrate

Situation Practical choice
An existing job has stable selectors and a controlled target Keep it temporarily, add explicit readiness checks, and monitor failures.
A new project needs modern JavaScript compatibility Choose a maintained browser automation library or a rendering API.
You need screenshots, PDFs, device views, or repeatable capture options Use an API such as ScreenshotNeo and keep the request logic in your application.
You need an AI agent to perform captures Use ScreenshotNeo’s MCP tools.

9. FAQ

Does page.open wait for every AJAX request?

No. Its documented callback reports the page load status. Your scraper must identify the application state that means the required data is ready.

Can I return a DOM element from page.evaluate?

No. Return serializable fields copied from that element, such as text, attributes, or numeric values.

Why does PhantomJS miss content that appears after a click?

The click may trigger asynchronous rendering. Perform the interaction, wait for the site’s resulting ready state, then evaluate the updated DOM.

Is PhantomJS suitable for a new production scraper?

Usually not. Development is suspended, the repository is archived, and the 2.x line is documented as deprecated and unmaintained. Treat it as legacy software and plan a migration.

Can ScreenshotNeo return HTML instead of an image?

ScreenshotNeo is designed for screenshots and PDFs. Its get_page_info MCP tool can provide page information, while the screenshot endpoint returns the requested capture format.