ScreenshotNeo

BlogGuides

What Is Cheerio in JavaScript?

Cheerio parses HTML and XML with a jQuery-like API. Learn what it does, what it cannot do, and when to use a real browser.

By the ScreenshotNeo team1 October 20266 min read

Cheerio is a JavaScript library for parsing HTML and XML into a queryable, jQuery-like structure. You load markup, select nodes with CSS selectors, read or modify them, and serialize the result. Cheerio does not open a visual browser, apply CSS, or execute page JavaScript, so it cannot see content that a single-page app inserts after load. See the official introduction.

Install Cheerio

In a Node.js project, install the package with npm:

npm install cheerio

Cheerio supports ES modules and CommonJS:

import * as cheerio from 'cheerio';
// CommonJS:
// const cheerio = require('cheerio');

First example: select and read HTML

import * as cheerio from 'cheerio';
const markup = '<h2 class="title">Hello world</h2>';
const $ = cheerio.load(markup);
console.log($('h2.title').text()); // Hello world
console.log($.html());             // serialized document

cheerio.load creates the API bound to a document. The $ function accepts CSS selectors; methods such as text, attr, find, and each operate on matched nodes.

Common extraction patterns

Read text and attributes

const title = $('article h1').first().text().trim();
const links = $('a[href]').map((_, el) => ({
  text: $(el).text().trim(), href: $(el).attr('href')
})).get();

Extract records

const products = $('.product').map((_, el) => {
  const card = $(el);
  return { name: card.find('.name').text().trim(),
    price: card.find('.price').text().trim(),
    url: card.find('a').attr('href') ?? null };
}).get();

Handle missing nodes

const description = $('meta[name="description"]').attr('content')?.trim() ?? '';
const image = $('.hero img').attr('src') ?? null;

Traverse relative to a node

$('li.item').each((index, element) => {
  const item = $(element);
  console.log(index, item.find('.label').text().trim());
});
const next = $('h2').first().next().text().trim();

Transform markup and serialize it

const $ = cheerio.load('<ul><li>One</li><li>Two</li></ul>');
$('li').addClass('item');
$('ul').prepend('<li class="item new">Zero</li>');
$('li').first().text('Updated');
$('li').last().remove();
console.log($.html());

Use $.html() for the serialized document, $.html(node) for a node, or $(selector).toString() for a matched fragment.

Load HTML, bytes, streams, or a URL

The loading guide documents load for strings, loadBuffer for raw bytes, stringStream for decoded text streams, decodeStream for raw byte streams, and fromURL for fetching a URL. Byte methods sniff encoding. fromURL refuses responses whose content type is neither HTML nor XML.

Load a local file

import { readFile } from 'node:fs/promises';
import * as cheerio from 'cheerio';
const html = await readFile('page.html', 'utf8');
const $ = cheerio.load(html);
console.log($('title').text().trim());

Fetch a URL

import * as cheerio from 'cheerio';
const $ = await cheerio.fromURL('https://example.com/');
console.log($('title').text().trim());

For production fetching, use timeouts, validate status codes, and follow the site’s terms and robots policy. A login page, bot check, PDF, or empty response is simply what Cheerio will parse; it will not become the browser-rendered page.

Parser choices: parse5 and htmlparser2

Cheerio uses parse5 by default for HTML, following HTML parsing rules to produce a browser-like tree. htmlparser2 is the default for XML and may also be selected for HTML when its more forgiving malformed-markup handling and lower memory use fit the input. These are the project’s qualitative descriptions, not benchmark results.

import * as cheerio from 'cheerio';
const $ = cheerio.load('<root><item>value</item></root>', { xml: true });
console.log($('item').text());

Choose parse5 when standards-oriented HTML correction matters. Choose XML/htmlparser2 mode for XML syntax or when malformed input makes its behavior preferable. Parser choice can change the resulting tree, so validate selectors against representative documents.

What Cheerio cannot do

  • No JavaScript execution: client-inserted content is absent from the original response.
  • No visual rendering: CSS layout, fonts, screenshots, canvas output, and computed styles are outside its scope.
  • No browser interaction: it cannot click, scroll, submit forms, or maintain browser storage like a real page.
  • No automatic resource loading: script and image tags do not cause their resources to be fetched.

When browser execution is needed, Cheerio’s introduction points to Puppeteer or Playwright; it also names jsdom for DOM emulation. A useful workflow is to render with a browser tool, then pass the resulting HTML to Cheerio for extraction.

Complete runnable scraper

Save as extract.mjs, then run node extract.mjs https://example.com/.

import * as cheerio from 'cheerio';
const target = process.argv[2];
if (!target) throw new Error('Usage: node extract.mjs <url>');
const response = await fetch(target, {
  headers: { 'user-agent': 'cheerio-example/1.0' },
  signal: AbortSignal.timeout(15000)
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const contentType = response.headers.get('content-type') || '';
if (!/html|xml/i.test(contentType)) throw new Error(`Expected HTML or XML, got ${contentType || 'unknown content type'}`);
const $ = cheerio.load(await response.text());
const result = {
  title: $('title').first().text().trim(),
  headings: $('h1, h2, h3').map((_, el) => $(el).text().trim()).get(),
  links: $('a[href]').map((_, el) => ({ text: $(el).text().trim(), href: $(el).attr('href') })).get()
};
console.log(JSON.stringify(result, null, 2));

Cheerio versus a browser tool

Requirement Cheerio Browser automation
Parse response HTML Yes Yes
CSS selection and traversal Yes Yes
Run page JavaScript No Yes
Layout, screenshots, canvas, computed styles No Yes
Lightweight server-side transformation Yes Usually excessive

Use Cheerio when markup is available and the job is extraction or transformation. Use Puppeteer or Playwright when the page must execute scripts or behave like a browser. Use jsdom when DOM emulation is the requirement.

Troubleshooting

Symptom Likely cause Fix
Selector is empty Content is client-injected, selector is wrong, or markup differs. Inspect the raw response, verify selector, or render with Puppeteer/Playwright.
Text is missing Value may be in an attribute or form property. Read attr('content'), attr('value'), or the correct element.
fromURL rejects response Content-Type is not HTML/XML, such as PDF or JSON. Check headers and use a format-specific parser or validate a manual fetch.
Unexpected nesting parse5 applies HTML tree-construction rules. Consider XML/htmlparser2 mode where suitable and adjust selectors to the parsed tree.
Garbled characters Bytes were decoded with the wrong encoding. Use loadBuffer or decodeStream for encoding sniffing.
CAPTCHA or login page Server returned an anti-bot or authenticated response. Respect access controls and use an approved authentication/browser flow.

Performance, reliability, and cost

  • Performance: parse once, keep selectors specific, extract only needed fields, and reuse the loaded object.
  • Memory: avoid retaining full documents across large batches; emit compact records and release references. Use htmlparser2 when its documented trade-offs fit.
  • Reliability: set network timeouts, check status and content type, handle redirects/retries deliberately, and record each source URL.
  • Cost: Cheerio is an npm dependency. Operational costs are network, compute, storage, and any browser service used before parsing; Cheerio provides no hosted scraping quota.

Or skip the browser setup

If you need a clean screenshot rather than HTML extraction, ScreenshotNeo provides a single API request. It accepts cookie and consent banners, removes 60+ known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.

It supports full-page or element captures, dark mode, device presets, custom viewports and retina scale, PDF output, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs, bulk capture, usage reporting, and an MCP server for AI agents. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The MCP server gives AI agents screenshot tools. Free usage is 1,000 screenshots monthly with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Is Cheerio a browser?

No. It parses supplied markup and exposes a jQuery-like API; it does not render or run JavaScript.

Can Cheerio scrape a React site?

Only the content in the response. Render first with a browser tool when the data appears after client execution.

Does Cheerio support XML?

Yes. XML mode uses htmlparser2 behavior.

Can I use Cheerio in the browser?

Cheerio is chiefly used in Node.js for server-side parsing. For browser DOM work, native DOM APIs or a browser-focused library are generally more suitable.