ScreenshotNeo

BlogComparisons

HTML Extraction APIs for Fully Rendered Web Pages

Compare rendered-HTML and structured extraction APIs, then build a reliable JavaScript rendering pipeline with waits, selectors, retries and cost controls.

By the ScreenshotNeo team1 October 20268 min read

Short answer: choose an API that runs a real browser when the content appears only after JavaScript executes. Use a rendered-HTML endpoint when your own parser needs the document, a selector-based endpoint when you already know the fields, and text or Markdown output when downstream processing does not need the DOM. Browserless documents /content for fully rendered HTML and /scrape for CSS-selector JSON. ScrapingBee documents JavaScript rendering for HTML, text, Markdown, screenshots and extraction rules. Crawl4AI documents both self-hosting and a hosted API.

No vendor documentation proves universal extraction accuracy or reliability. Test the pages, fields, latency, concurrency and cost that matter to your workload.

What “fully rendered HTML” means

A normal HTTP client receives the server response. On a JavaScript-heavy site, that response may contain almost no article text: a root element, script tags and configuration data. A rendering API starts a browser, loads the page, executes JavaScript, waits for a condition, and then returns the resulting DOM or selected data.

Rendering does not guarantee that every page is accessible or that an extraction is correct. Bot checks, login walls, consent dialogs, infinite scrolling, shadow DOM and client-side errors can still change the result.

Pick the output before picking the API

Need Output Typical choice
Your parser needs the document structure Rendered HTML Browserless /content or ScrapingBee HTML output
You know the fields and selectors Structured JSON Browserless /scrape with CSS selectors
You need readable content for indexing or an LLM Text or Markdown ScrapingBee text/Markdown output, or render HTML and clean it yourself
You need a screenshot or PDF as evidence Image/PDF A browser screenshot/PDF endpoint such as Browserless, or ScreenshotNeo

How to evaluate an API

  1. Inspect the first response. Fetch the URL without JavaScript. If the required text is already present, a browser may be unnecessary.
  2. Define the readiness condition. Prefer a selector, network-idle condition or application event tied to the content. A fixed sleep is a fallback.
  3. Choose the extraction boundary. Return the whole DOM only when your parser needs it. Selector extraction reduces payload and post-processing.
  4. Test representative pages. Include slow pages, pages with consent banners, pagination, lazy content, authentication and occasional failures.
  5. Measure your own workload. Record field completeness, error classes, p50/p95 latency, concurrency, bytes returned and credits or request cost.

DIY: render and extract with Playwright

When you need complete browser control, run Playwright yourself. The following Node.js program waits for an article selector, removes obvious non-content elements, and writes rendered HTML.

import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';

const url = process.argv[2] ?? 'https://example.com';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });

try {
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
  await page.waitForSelector('article, main, [role="main"]', { timeout: 20000 });
  await page.evaluate(() => {
    document.querySelectorAll('script, style, noscript, nav, footer, [aria-label*="chat" i]').forEach(el => el.remove());
  });
  const html = await page.content();
  await writeFile('rendered.html', html, 'utf8');
  console.log('Wrote rendered.html');
} finally {
  await browser.close();
}

Install and run it with:

npm init -y
npm install playwright
npx playwright install chromium
node render.mjs https://your-site.example/page

Python Playwright version

from pathlib import Path
from playwright.sync_api import sync_playwright

url = "https://example.com"
with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1440, "height": 900})
    try:
        page.goto(url, wait_until="domcontentloaded", timeout=45_000)
        page.wait_for_selector("article, main, [role='main']", timeout=20_000)
        page.locator("script, style, noscript, nav, footer").evaluate_all(
            "els => els.forEach(el => el.remove())"
        )
        Path("rendered.html").write_text(page.content(), encoding="utf-8")
    finally:
        browser.close()
pip install playwright
playwright install chromium
python render.py

Extract known fields in the browser

const data = await page.locator('article').evaluate(article => ({
  title: article.querySelector('h1')?.textContent?.trim() ?? null,
  paragraphs: [...article.querySelectorAll('p')].map(p => p.textContent.trim()),
  links: [...article.querySelectorAll('a[href]')].map(a => ({
    text: a.textContent.trim(), href: a.href
  }))
}));
console.log(JSON.stringify(data, null, 2));

Hosted API patterns

ScrapingBee HTML API

ScrapingBee documents JavaScript rendering enabled by default for its HTML API and describes support for React, Angular, JQuery and Vue applications. Its documentation also covers HTML, text, Markdown, screenshots, extraction rules, waits and proxy configuration. Credit usage varies by configuration: the documentation lists classic proxy without JavaScript at 1 credit, classic proxy with JavaScript at 5, premium proxy without JavaScript at 10, premium proxy with JavaScript at 25, stealth proxy with JavaScript at 75, and AI extraction at an additional 5 credits. Confirm current terms before budgeting.

Use a wait tied to the content you require. For example, wait for the article selector rather than sleeping for an arbitrary number of seconds. Keep proxy and stealth settings proportional to the target: they add cost and do not fix an application-level selector bug.

Browserless REST APIs

Browserless documents separate REST endpoints: /content returns fully rendered HTML, /scrape extracts JSON with CSS selectors, and /smart-scrape provides a cascading approach for blocked or JavaScript-heavy sites. It also documents screenshot and other browser-task endpoints. The vendor describes these as HTTP endpoints for screenshots, PDFs, content scraping, file downloads, function execution and website unblocking.

A generic request shape looks like this; use the endpoint and authentication format from your Browserless account documentation:

curl -X POST 'https://YOUR_BROWSERLESS_HOST/content?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' \
  --data '{"url":"https://example.com","waitFor":"article"}'

For selector extraction, send the target URL and selectors to the documented /scrape endpoint and validate that every required field is present before accepting the record.

Crawl4AI

Crawl4AI presents an open-source crawler that can be self-hosted and also documents a hosted API for scraping, search and extraction. The surfaced documentation identifies itself as version 0.9.x; verify current hosted availability, API shape and pricing before committing to it.

Or skip the browser setup

If your deliverable is a screenshot or PDF of the rendered page, ScreenshotNeo provides a single request instead of maintaining Playwright workers. Its API accepts a URL and returns PNG, JPEG, WebP or PDF; options cover full-page capture with lazy images, CSS-selector element capture, device presets, viewport and retina scale, dark mode, waits, custom CSS and JavaScript, click and hide selectors, blocked resources, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks and bulk capture.

See the ScreenshotNeo API documentation for the current parameter list.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed; response headers identify the page verdict and whether the request was billed. ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account and start with the included 1,000 monthly screenshots.

Rendering controls that affect extraction

Waits

  • Selector wait: best when a specific element signals usable content.
  • Network idle: useful for applications that finish loading through fetch/XHR, but analytics or long polling can prevent it.
  • Fixed delay: predictable but either wastes time or races the application.
  • Application event: most precise when you control the page and can expose a ready flag.

Lazy loading and infinite scroll

Scroll incrementally, wait after each scroll, and stop when the document height no longer grows or a next-page control disappears. Set a maximum scroll count so a broken page cannot run forever.

Accept or dismiss consent only when your legal and data requirements allow it. Record the action in metadata so an extraction can be reproduced. Remove overlays after the page has reached the required state; removing them too early can prevent the application from initializing.

Authentication and regional content

Use an isolated browser context per account. Supply cookies or headers deliberately, avoid logging secrets, and set timezone or geolocation when the page varies by locale. Verify that the returned content belongs to the intended account and region.

Reliability checklist

  • Use a bounded navigation timeout and a separate extraction timeout.
  • Retry transient network failures with exponential backoff and jitter.
  • Do not retry deterministic selector failures indefinitely.
  • Store the final URL, HTTP status, title, extraction version and failure reason.
  • Validate required fields and minimum text length before writing a record.
  • Capture a diagnostic screenshot or HTML sample for failed jobs when policy permits.
  • Limit concurrency to the provider plan and your own CPU and memory capacity.
  • Cache pages whose freshness requirements allow it.

Performance and cost

Browser startup, JavaScript execution, proxy routing and large assets dominate latency. Reuse browser processes, block fonts, ads and analytics when they are irrelevant, wait for the smallest useful selector, and return only the fields you need. Measure p95 latency rather than relying on a single successful request.

Hosted pricing models differ. ScrapingBee publishes monthly plans and concurrency: the surfaced pricing page lists Hobby at $19/month for 75,000 credits and 25 concurrent requests; Freelance at $49 for 250,000 credits and 50 concurrency; Startup at $99 for 1,000,000 credits and 100 concurrency; Business at $249 for 3,000,000 credits and 200 concurrency; and Business+ at $599 for 8,000,000 credits and 400 concurrency. It also advertises 1,000 free credits. These are vendor figures accessed in 2026 and can change.

For self-hosting, include browser CPU, memory, patching, queueing, proxy and observability costs. For every option, calculate cost per successfully complete record, not cost per request, because retries and partial extractions change the real total.

Troubleshooting

Symptom Likely cause Fix
HTML contains only a root element JavaScript did not run or the request used a plain HTTP client Enable browser rendering and wait for the content selector
Selector returns zero nodes Wrong selector, iframe, shadow DOM or content not ready Inspect the rendered DOM, handle frames/shadow roots, and change the wait condition
Fields are intermittently missing Race condition, slow API call or A/B variant Wait for a field-specific signal, retry transient failures, and validate required fields
Infinite scroll never finishes Network idle never occurs or the page continuously loads Use bounded scrolling and a stop condition based on document height or item count
Content is a bot-check page Target blocks automated browsers Respect site rules, use documented proxy options where appropriate, and classify the result as a failure
Unexpected language or prices Locale, timezone, cookies or geolocation differ Set these values explicitly and test each target region
High cost Premium proxy, stealth mode, AI extraction or repeated rendering Use the least expensive mode that meets the requirement, cache results and avoid duplicate jobs

FAQ

Is rendered HTML the same as the original source?

No. It is the DOM after browser scripts and mutations have run. Keep the original response too when provenance matters.

Should I extract HTML or JSON?

Choose HTML when schemas change or your parser needs context. Choose selector JSON when fields are stable and you want smaller, validated output.

Can an API guarantee complete page data?

No. Readiness, access controls, personalization and application bugs can all affect completeness. Validate required fields.

Do I need a browser for every URL?

No. First check whether the initial HTTP response already contains the required content. Render only the pages that need JavaScript.

How should I compare providers?

Run the same representative URL set and compare completeness, latency, failure recovery, concurrency, geographic behavior, operational effort and cost per successful record. The available research contains no independent like-for-like benchmark.