ScreenshotNeo

BlogComparisons

JavaScript Web Scraping Libraries: Features and Limitations

Compare Cheerio, Puppeteer, Playwright and Crawlee, with runnable Node.js examples, production guidance, troubleshooting and compliance notes.

By the ScreenshotNeo team29 September 20269 min read

JavaScript Web Scraping Libraries: Features and Limitations

Direct answer: start with Cheerio when the fields you need are in the initial HTML response. Move to Playwright when JavaScript builds the page, you need clicks or form input, or browser behavior matters. Choose Puppeteer for Chrome or Firefox automation when cross-engine WebKit support is unnecessary. Choose Crawlee when you need queues, persistence, retries, proxies, sessions, routing, scaling, or a crawler that can switch between HTTP and browsers.

A practical production design is tiered: fetch and parse with Cheerio first, then send only JavaScript-dependent URLs to Playwright or Puppeteer. This keeps resource use low while preserving a browser path for single-page apps (SPAs), protected flows, screenshots, and interactive pages.

What each JavaScript scraping library actually does

Library Best fit Strengths Limitations
Cheerio Static HTML/XML Very low overhead; jQuery-like selectors and traversal No visual rendering, external resources, or JavaScript execution
Puppeteer Chrome/Firefox automation, screenshots, PDFs and UI workflows High-level API over CDP/WebDriver BiDi; headless by default Heavier runtime; browser install scripts can be blocked
Playwright Cross-browser interaction and robust synchronization Chromium, Firefox, WebKit, Chrome and Edge; locators and auto-waiting Matching browser binaries must be installed and maintained
Crawlee Production crawlers Queues, storage, retries, proxies, sessions, routing and scaling More dependencies; browser crawlers are installed separately

Cheerio’s documentation is explicit: “Cheerio is not a web browser.” It parses markup without CSS, visual rendering, external-resource loading, or JavaScript execution. That is why it is fast and why client-rendered content can be missing. Read the Cheerio introduction.

Choose by the page’s rendering model

Use Cheerio for server-rendered HTML

Use Cheerio when a normal HTTP response contains the title, links, prices, article body, or other fields. It is usually the simplest and least expensive option. Inspect the response body first; if the values are present, a browser adds startup and memory cost without adding data.

Choose a parser when the data is in the response; use a browser when JavaScript must render it.
Choose a parser when the data is in the response; use a browser when JavaScript must render it.
npm install cheerio
import * as cheerio from 'cheerio';

const response = await fetch('https://example.com/products');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const $ = cheerio.load(html);

const products = $('.product').map((_, el) => ({
  name: $(el).find('.name').text().trim(),
  price: $(el).find('.price').text().trim(),
  href: new URL($(el).find('a').attr('href'), 'https://example.com').href
})).get();

console.log(products);

Cheerio cannot click a button, execute a script, wait for an XHR, render a canvas, or see content that arrives after hydration. A common SPA symptom is an almost empty root element in the downloaded HTML.

Use Playwright for JavaScript-rendered pages

Playwright controls Chromium, Firefox, WebKit, Chrome, and Edge. Its locators auto-wait for actionability, which reduces hand-written sleeps and race conditions. Install the package and the browser binaries:

npm install playwright
npx playwright install
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
await page.locator('[data-testid="load-more"]').click();
await page.locator('.product').first().waitFor();
const products = await page.locator('.product').evaluateAll(nodes =>
  nodes.map(node => ({
    name: node.querySelector('.name')?.textContent?.trim(),
    price: node.querySelector('.price')?.textContent?.trim()
  }))
);
await browser.close();
console.log(products);

Prefer a meaningful readiness condition such as a locator, a response, or a page-specific state. networkidle can be unsuitable for pages with analytics, polling, or long-lived connections.

Use Puppeteer for Chrome-oriented automation

Puppeteer runs headless by default and is a good fit for Chrome workflows, screenshots, PDFs, and browser state. Install it with its browser download scripts enabled:

npm install puppeteer
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product');
const products = await page.$$eval('.product', nodes => nodes.map(node => ({
  name: node.querySelector('.name')?.textContent?.trim(),
  price: node.querySelector('.price')?.textContent?.trim()
})));
await browser.close();
console.log(products);

If your package manager blocks install scripts, Puppeteer may not download a browser and will fail at runtime. Use an approved browser path or allow the install step in your build process.

Use Crawlee when crawling is the system

Crawlee provides CheerioCrawler, PuppeteerCrawler, and PlaywrightCrawler behind common crawler concepts. It adds persistent request queues, storage, retries, routing, proxy rotation, sessions, Docker support, and resource-based scaling. Install the core package plus the crawler you need:

npm install crawlee playwright
npx playwright install chromium
import { PlaywrightCrawler } from 'crawlee';

const crawler = new PlaywrightCrawler({
  maxRequestsPerCrawl: 100,
  requestHandler: async ({ page, request, log }) => {
    await page.locator('.product').first().waitFor();
    const products = await page.locator('.product').evaluateAll(nodes =>
      nodes.map(node => ({
        url: request.url,
        name: node.querySelector('.name')?.textContent?.trim(),
        price: node.querySelector('.price')?.textContent?.trim()
      }))
    );
    log.info(`Found ${products.length} products`);
    console.log(products);
  }
});

await crawler.run(['https://example.com/catalog']);

Crawlee’s CheerioCrawler is efficient but cannot handle JavaScript rendering. Its PuppeteerCrawler and PlaywrightCrawler use a headless browser. The default install does not bundle those browser integrations, so add the package and browser binaries explicitly.

Complete browser options that affect scraping

  • URL and method: follow redirects deliberately and check final URLs.
  • Timeouts: set navigation, locator, and overall job limits; avoid unlimited waits.
  • Readiness: wait for a selector, a specific response, a URL change, or an application state.
  • Frames and tabs: inspect if content is inside an iframe or opened in a new page.
  • Retries: retry transient network failures with backoff, but cap attempts for deterministic 4xx responses.

Browser context and identity

Use a fresh context for isolation. Set viewport, locale, timezone, user agent, permissions, cookies, extra headers, HTTP authentication, and geolocation only when your use case requires them. Keep credentials in a secret manager; never hard-code production tokens in source.

Extraction choices

Use stable attributes such as data-testid or semantic roles instead of brittle generated class names. Normalize whitespace, currencies, dates, and relative URLs. Store the source URL and capture timestamp with every record so downstream users can audit a result.

Resource control

Block images, fonts, ads, trackers, or third-party requests only when they are irrelevant to the data. Blocking an API request that supplies the page will produce an empty result. Limit concurrency to what the target and your host can handle, and reuse a browser process while creating isolated contexts.

Scraping SPAs and data loaded after hydration

  1. Fetch the URL with Cheerio and inspect whether the required fields exist.
  2. If they do not, open the page in Playwright or Puppeteer.
  3. Wait for a field-specific locator or the API response that populates it.
  4. Extract rendered DOM or parse the JSON response when that endpoint is authorized and stable.
  5. Record missing fields and diagnostics instead of silently returning an empty object.

Many SPAs expose data through XHR or fetch calls. Browser network logging can reveal the request, but do not bypass authentication, rate limits, or access controls. A browser is also necessary when the value depends on client-side computation, scrolling, clicking, or a visual state.

Do-it-yourself screenshots and PDFs

For a local screenshot, Playwright can capture the viewport or the full page:

import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage({ viewport: { width: 1365, height: 768 }, deviceScaleFactor: 2 });
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await page.screenshot({ path: 'page.png', fullPage: true });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
await browser.close();

In CI, install the exact browser version expected by your package, use deterministic fonts where possible, and wait for images or application data before capture. Full-page screenshots can be large; resize or compress them before storing or sending them.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page and element capture, dark mode, device presets, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture, usage, and the OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
print(r.headers.get("X-Page-Verdict"), r.headers.get("X-Billed"))
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Performance, reliability, and cost

Concern Cheerio Browser tools Crawlee
Startup and memory Lowest Highest Depends on crawler type
JavaScript execution No Yes Yes with browser crawlers
Scaling controls Build yourself Build yourself Queues, retries and scaling included
Operational work HTTP and parser only Browser binaries and sandboxing Framework plus browser dependencies

Measure your own workload; no neutral cross-library benchmark is established by the supplied research. Reduce cost by filtering URLs, caching immutable pages, using Cheerio for the first attempt, limiting browser concurrency, and closing contexts. For reliability, pin library versions, install matching browser binaries during deployment, log navigation status and final URL, save a small diagnostic artifact for failures, and make jobs idempotent.

Troubleshooting common failures

Symptom Likely cause Fix
Empty fields with Cheerio Content is client-rendered Use Playwright/Puppeteer or an authorized data endpoint
Browser executable not found Install script was blocked or binaries are missing Run the Playwright install command or permit Puppeteer’s browser download
Timeout waiting for selector Wrong selector, consent wall, slow API, or iframe Inspect the DOM, handle consent, wait for the correct frame or response, and set a bounded timeout
Flaky results Fixed sleeps or changing page state Use locators, response waits, stable attributes, and isolated contexts
Works locally, fails in CI Missing dependencies, sandbox restrictions, fonts, or viewport differences Pin versions, install browsers in the image, configure sandboxing, and standardize the environment
429 or blocked requests Rate limits or site defenses Slow down, respect published limits, cache results, and obtain permission
Screenshot contains overlays Cookie banner, popup, or chat widget Dismiss it with a locator or use ScreenshotNeo’s cleanup options
Consent banners and overlays can be handled before a screenshot is saved.
Consent banners and overlays can be handled before a screenshot is saved.

Compliance checklist

RFC 9309 defines robots.txt processing as a requested protocol and states that its rules are not access authorization. Treat robots.txt as one input, then review terms of service, authentication boundaries, privacy obligations, copyright, rate limits, and applicable law. Cache robots.txt conservatively; the RFC says a cached file generally should not be used for more than 24 hours unless it is unreachable.

  • Identify the data owner and your lawful purpose.
  • Honor robots.txt, terms, authentication, and rate limits.
  • Minimize personal data and protect credentials.
  • Provide deletion and retention controls where required.
  • Keep an audit trail of URLs, timestamps, and failures.

FAQ

Can Cheerio scrape a React or Vue site?

Only if the required data is already present in the initial HTML. Otherwise it cannot execute the JavaScript that renders the page.

Is Playwright always better than Puppeteer?

No. Playwright is the stronger choice for cross-browser coverage and auto-waiting. Puppeteer remains appropriate for Chrome or Firefox workflows when its API and ecosystem fit the project.

When should I introduce Crawlee?

Introduce it when scheduling, persistence, retries, sessions, proxies, routing, or scaling become application requirements rather than one-off script details.

Should I scrape an internal API instead of the DOM?

It can be simpler and more stable when the endpoint is authorized and intended for your use. Keep the same compliance, authentication, and rate-limit review.

Do browser libraries include browsers?

They depend on browser binaries. Playwright may require a fresh install after an update, and Puppeteer can fail when package-manager install scripts are blocked.