ScreenshotNeo

BlogHow-to

Data Extraction in Node.js: Cheerio, jsdom, Playwright, and Streaming

Choose the right Node.js extraction approach: stream and validate responses, parse static markup with Cheerio, or use jsdom and Playwright when browser behavior matters.

By the ScreenshotNeo team29 September 202610 min read

Data Extraction in Node.js: Cheerio, jsdom, Playwright, and Streaming

For a static page, fetch its HTML and parse it with Cheerio. Use jsdom when your extraction code needs DOM APIs, and use Playwright when the data appears only after browser JavaScript runs or depends on browser network behavior. For large responses, stream bytes with backpressure, check status and content type, and validate extracted fields before saving them.

This guide builds a small static-page extractor, then covers streaming, encoding, DOM emulation, browser rendering, failures, and operational concerns. Respect a site’s terms, access controls, and applicable robots guidance.

1. Define what counts as a successful extraction

Before choosing a library, write down the source contract. Record the URL or API endpoint, expected content type, fields, pagination scheme, authentication method, and any rate limits. Decide which fields are required and what validation makes each value usable.

Keep provenance with every record: at minimum, the source URL and retrieval time. Normalize whitespace, URLs, dates, and numbers at the boundary. A missing required field should be an observable extraction failure, not a silently emitted partial record.

  1. Identify one representative page and inspect the response or browser-rendered page.
  2. Check whether the needed values exist in the delivered HTML.
  3. Choose a parser based on where the data comes from.
  4. Validate a few records and add fixtures so layout changes become visible.

2. Fetch and parse static HTML with Cheerio

Cheerio is a good fit when the server already delivered the markup containing the fields. It parses HTML or XML and offers jQuery-like traversal. It does not render a browser, load external resources, or execute page JavaScript. If a client-side app inserts the target data only after execution, parsing the initial response will not find it. See the [Cheerio introduction](https://cheerio.js.org/docs/intro/) and [loading guide](https://cheerio.js.org/docs/basics/loading/).

Install Cheerio in a Node project with npm install cheerio. This runnable example uses its URL loader and checks each record rather than assuming selectors always match:

import * as cheerio from 'cheerio';

const url = 'https://example.com/catalog';
const $ = await cheerio.fromURL(url);
const retrievedAt = new Date().toISOString();
const records = [];

$('.product').each((_, element) => {
  const title = $(element).find('.product-title').text().trim();
  const href = $(element).find('a').attr('href');
  const priceText = $(element).find('.price').text().trim();

  if (!title || !href || !priceText) return;
  records.push({
    title,
    url: new URL(href, url).href,
    priceText,
    sourceUrl: url,
    retrievedAt,
  });
});

if (records.length === 0) {
  throw new Error(`No complete product records found at ${url}`);
}
console.log(JSON.stringify(records, null, 2));

Replace the example URL and selectors with selectors from the source you are authorized to access. Relative links are resolved against the page URL. Keep the original text for values such as prices until you have a deliberate locale-aware conversion rule.

Cheerio offers several loading paths. load() accepts a string; loadBuffer() accepts bytes and detects encoding; stringStream() and decodeStream() support streaming input; and fromURL() fetches a URL. Its URL loader follows up to five redirects, rejects non-2xx responses and non-markup content types, and uses the final URL as the base URI. If you pass request options, specify the method; custom headers replace the default header set. Check the [loader documentation](https://cheerio.js.org/docs/basics/loading/) before customizing requests.

By default, Cheerio uses standards-oriented parse5 for HTML and htmlparser2 for XML. Its documentation describes htmlparser2 as faster, lower-memory, and more forgiving of malformed markup. Treat parser choice as an input-quality and workload decision; verify output against representative fixtures. More details are in [Cheerio’s parser configuration guide](https://cheerio.js.org/docs/advanced/configuring-cheerio/).

3. Handle large responses without unbounded buffering

Node’s HTTP API is deliberately low-level and does not buffer entire requests or responses, which supports streaming large or chunk-encoded messages. Use streams and backpressure when response size can grow beyond a comfortable in-memory bound. The [Node HTTP documentation](https://nodejs.org/api/http.html) describes this model. Node’s [Web Streams API](https://nodejs.org/api/webstreams.html) follows WHATWG streams and provides Readable.toWeb() and Readable.fromWeb() conversions between web and Node streams.

A reliable extraction pipeline checks the response before parsing and validates records afterward.
A reliable extraction pipeline checks the response before parsing and validates records afterward.

Streaming the network response does not automatically make every parser or transformation bounded-memory. A typical HTML extraction still needs a DOM-like representation, and a simple whole-document parse may retain the document. For genuinely large structured feeds, process incrementally with a streaming parser or consume an API that pages records. Use Cheerio’s stream loaders when they match the input and decoding needs, and measure memory with realistic documents.

For a known UTF-8 endpoint, the following pattern demonstrates Node’s HTTP stream and status checks. It bounds collected data and stops if the response exceeds the configured limit. The resulting string is then suitable for a normal parser:

import { get } from 'node:https';

const url = new URL('https://example.com/catalog');
const maxBytes = 5 * 1024 * 1024;

const html = await new Promise((resolve, reject) => {
  const req = get(url, {
    headers: { 'user-agent': 'CatalogExtractor/1.0' },
    timeout: 15000,
  }, (res) => {
    if (res.statusCode < 200 || res.statusCode >= 300) {
      res.resume();
      reject(new Error(`Unexpected HTTP status ${res.statusCode}`));
      return;
    }
    const type = String(res.headers['content-type'] || '').toLowerCase();
    if (!type.includes('text/html')) {
      res.resume();
      reject(new Error(`Expected HTML, received ${type || 'unknown type'}`));
      return;
    }
    let size = 0;
    const chunks = [];
    res.on('data', (chunk) => {
      size += chunk.length;
      if (size > maxBytes) {
        req.destroy(new Error('Response exceeded size limit'));
        return;
      }
      chunks.push(chunk);
    });
    res.on('end', () => resolve(Buffer.concat(chunks).toString('utf8')));
    res.on('error', reject);
  });
  req.on('timeout', () => req.destroy(new Error('Request timed out')));
  req.on('error', reject);
});

console.log(html.length);

This example caps the response but still concatenates it for parsing; it is not a constant-memory HTML extraction pipeline. For unknown encodings, retain bytes and use a byte-aware loader such as Cheerio’s loadBuffer() or decodeStream() rather than decoding blindly as UTF-8. For many concurrent requests, also cap concurrency and response sizes.

4. Decide between Cheerio, jsdom, and Playwright

Tool Execution model Use it when Tradeoff
Cheerio Parses delivered markup Required fields are already in HTML or XML Does not run site JavaScript or load browser resources
jsdom Emulates many DOM and HTML standards in JavaScript Extraction code expects document, selectors, or DOM-shaped behavior Not a full browser; does not guarantee all real-browser behavior
Playwright Runs a browser and exposes network controls Content depends on browser execution, navigation, or request behavior Browser setup and execution generally use more resources than static parsing

jsdom is a pure-JavaScript implementation of many WHATWG DOM and HTML standards. Its [README](https://github.com/jsdom/jsdom/blob/main/README.md) describes it as useful for emulating enough browser behavior to test and scrape web applications. Choose it when DOM semantics are the requirement; choose a real browser when behavior depends on the full browser environment.

Choose Cheerio, jsdom, or Playwright based on when and where the target data appears.
Choose Cheerio, jsdom, or Playwright based on when and where the target data appears.

Playwright is also useful when the network itself is part of the extraction problem. Its routing APIs can fetch a response for inspection or modification before fulfilling a route, change headers, and set a maximum redirect count. Request lifecycle events include request, response, requestfinished, and requestfailed. A 404 or 503 can still be a completed response, so inspect status explicitly. See the [route API](https://playwright.dev/docs/api/class-route) and [request API](https://playwright.dev/docs/api/class-request).

5. Extract browser-rendered data with Playwright

Use Playwright when the value is absent from the initial markup and appears after page scripts run. Install the package and browser for your environment using the [Playwright installation guide](https://playwright.dev/docs/intro). This example waits for a product selector, then checks that extracted fields are present:

import { chromium } from 'playwright';

const url = 'https://example.com/catalog';
const browser = await chromium.launch();
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
  await page.locator('.product').first().waitFor({ timeout: 10000 });
  const records = await page.locator('.product').evaluateAll((items) =>
    items.map((item) => ({
      title: item.querySelector('.product-title')?.textContent?.trim() || '',
      href: item.querySelector('a')?.getAttribute('href') || '',
      priceText: item.querySelector('.price')?.textContent?.trim() || '',
    }))
  );
  const complete = records.filter((r) => r.title && r.href && r.priceText);
  if (complete.length === 0) throw new Error('No complete records found');
  console.log(JSON.stringify(complete, null, 2));
} finally {
  await browser.close();
}

Pick the wait condition based on the page: waiting for a relevant selector is often clearer than sleeping for an arbitrary interval. A navigation’s document load does not necessarily mean a client app has finished loading its data. Set explicit navigation and selector timeouts, and close browser contexts even when extraction throws.

6. Configure requests and normalize records

Whatever retrieval path you choose, make the request behavior explicit. Set a reasonable timeout, a clear user agent, and only the headers and cookies required by the source. Bound redirects and verify the final URL if redirects can change the intended resource. Check status and content type before parsing; an HTML error page or login page can be syntactically valid HTML while being the wrong input.

For pagination, define a maximum page count and detect repeated next-page URLs. For authenticated sources, keep secrets out of logs and source code. Treat 401 and 403 responses as access failures, not prompts to bypass access controls. Retry only transient failures, with a bounded attempt count and delay; do not retry deterministic parse failures indefinitely.

Normalize output into a stable schema. Parse numbers and dates with the source’s locale and timezone in mind. Preserve the raw source value when conversion could be ambiguous. Store source URL, retrieval time, and a run identifier so a bad batch can be traced and safely retried.

7. Reliability, performance, and cost

Static parsing is usually the simplest and lightest option when the source HTML contains the fields. Streaming can reduce buffering for transport and large inputs, but parsing the whole document can still consume memory. DOM emulation adds browser-like APIs, while browser automation adds rendering and network activity; use those costs only when the data requires them.

Reliability comes from checks around the parser: request timeout, status, content type, redirect policy, response-size bound, expected record count or required fields, and structured error logs. Keep fixtures from known-good pages and re-run extraction against them when selectors change. For live runs, alert on sudden zero-record results or a high missing-field rate instead of treating them as successful empty exports.

There is no universal throughput number for these tools: page size, scripts, selectors, concurrency, and machine resources vary. Benchmark with representative inputs, monitor memory and latency, and tune concurrency conservatively. Persist checkpoints for paginated jobs so transient failures do not force a full restart. Minimize repeated fetches where permitted and useful, and set a bounded cache policy appropriate to how quickly the source changes.

8. Troubleshooting common failures

Symptom Likely cause Fix
Selectors find nothing Markup changed, wrong selector, or data is client-rendered Inspect the response HTML; validate selectors on a fixture; switch to Playwright if execution is required.
Cheerio rejects the URL response Non-2xx status or non-markup content type Log status and content type, verify the endpoint and redirect destination, and handle the error deliberately.
Text has replacement characters Bytes decoded using the wrong encoding Use a byte-aware loader such as loadBuffer() or decodeStream().
Only the first page is extracted Pagination was not modeled Follow the documented next link or page token with a maximum-page guard and duplicate detection.
Playwright times out waiting for data Wrong readiness condition, slow request, or page error Wait for the actual data selector, inspect console/network events, and set a bounded timeout appropriate to the source.
A 404 appears as a completed Playwright request HTTP error responses still produce a response lifecycle Read the response status and fail or branch on non-success codes.
Process memory grows during a run Responses, DOMs, pages, or concurrency are unbounded Cap input bytes and concurrency, close pages, avoid retaining documents, and stream or paginate where possible.
Some records are incomplete Optional markup, selector drift, or partial page state Validate required fields, report missing-field counts, and store failed records for inspection rather than silently discarding the problem.

9. Or skip the browser setup

If your goal is a screenshot of the page rather than a structured record, [ScreenshotNeo](https://screenshotneo.com) provides a website screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF. The API supports full-page capture with lazy images loaded, element capture by CSS selector, custom viewport and device presets, dark mode, custom CSS and JavaScript, wait conditions, request blocking, headers and cookies, and more. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed, and cache hits cost nothing; response headers report page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Sign up free for 1,000 screenshots a month with no card.

10. FAQ

Can Cheerio scrape a single-page application?

It can parse any data present in the delivered markup, including server-rendered parts of an app. It cannot execute the scripts that populate data only after the page loads. Use Playwright for that browser-dependent path.

When should I use an API instead of scraping HTML?

Use a documented endpoint when it provides the fields you need and you are authorized to call it. APIs usually give a more stable data shape than selectors tied to presentation markup.

Should I use jsdom or Playwright for selectors?

Use jsdom when your code needs DOM APIs against markup. Use Playwright when real browser execution, navigation, or network behavior determines what data exists.

Does streaming guarantee low memory use?

No. It avoids buffering the transport response wholesale, but later parsing or retaining records can still use substantial memory. Bound each stage and measure the complete pipeline.