What Is Cheerio in JavaScript?
Cheerio parses HTML and XML with a jQuery-like API. Learn what it does, what it cannot do, and when to use a real browser.
Cheerio is a JavaScript library for parsing HTML and XML into a queryable, jQuery-like structure. You load markup, select nodes with CSS selectors, read or modify them, and serialize the result. Cheerio does not open a visual browser, apply CSS, or execute page JavaScript, so it cannot see content that a single-page app inserts after load. See the official introduction.
Install Cheerio
In a Node.js project, install the package with npm:
npm install cheerio
Cheerio supports ES modules and CommonJS:
import * as cheerio from 'cheerio';
// CommonJS:
// const cheerio = require('cheerio');
First example: select and read HTML
import * as cheerio from 'cheerio';
const markup = '<h2 class="title">Hello world</h2>';
const $ = cheerio.load(markup);
console.log($('h2.title').text()); // Hello world
console.log($.html()); // serialized document
cheerio.load creates the API bound to a document. The $ function accepts CSS selectors; methods such as text, attr, find, and each operate on matched nodes.
Common extraction patterns
Read text and attributes
const title = $('article h1').first().text().trim();
const links = $('a[href]').map((_, el) => ({
text: $(el).text().trim(), href: $(el).attr('href')
})).get();
Extract records
const products = $('.product').map((_, el) => {
const card = $(el);
return { name: card.find('.name').text().trim(),
price: card.find('.price').text().trim(),
url: card.find('a').attr('href') ?? null };
}).get();
Handle missing nodes
const description = $('meta[name="description"]').attr('content')?.trim() ?? '';
const image = $('.hero img').attr('src') ?? null;
Traverse relative to a node
$('li.item').each((index, element) => {
const item = $(element);
console.log(index, item.find('.label').text().trim());
});
const next = $('h2').first().next().text().trim();
Transform markup and serialize it
const $ = cheerio.load('<ul><li>One</li><li>Two</li></ul>');
$('li').addClass('item');
$('ul').prepend('<li class="item new">Zero</li>');
$('li').first().text('Updated');
$('li').last().remove();
console.log($.html());
Use $.html() for the serialized document, $.html(node) for a node, or $(selector).toString() for a matched fragment.
Load HTML, bytes, streams, or a URL
The loading guide documents load for strings, loadBuffer for raw bytes, stringStream for decoded text streams, decodeStream for raw byte streams, and fromURL for fetching a URL. Byte methods sniff encoding. fromURL refuses responses whose content type is neither HTML nor XML.
Load a local file
import { readFile } from 'node:fs/promises';
import * as cheerio from 'cheerio';
const html = await readFile('page.html', 'utf8');
const $ = cheerio.load(html);
console.log($('title').text().trim());
Fetch a URL
import * as cheerio from 'cheerio';
const $ = await cheerio.fromURL('https://example.com/');
console.log($('title').text().trim());
For production fetching, use timeouts, validate status codes, and follow the site’s terms and robots policy. A login page, bot check, PDF, or empty response is simply what Cheerio will parse; it will not become the browser-rendered page.
Parser choices: parse5 and htmlparser2
Cheerio uses parse5 by default for HTML, following HTML parsing rules to produce a browser-like tree. htmlparser2 is the default for XML and may also be selected for HTML when its more forgiving malformed-markup handling and lower memory use fit the input. These are the project’s qualitative descriptions, not benchmark results.
import * as cheerio from 'cheerio';
const $ = cheerio.load('<root><item>value</item></root>', { xml: true });
console.log($('item').text());
Choose parse5 when standards-oriented HTML correction matters. Choose XML/htmlparser2 mode for XML syntax or when malformed input makes its behavior preferable. Parser choice can change the resulting tree, so validate selectors against representative documents.
What Cheerio cannot do
- No JavaScript execution: client-inserted content is absent from the original response.
- No visual rendering: CSS layout, fonts, screenshots, canvas output, and computed styles are outside its scope.
- No browser interaction: it cannot click, scroll, submit forms, or maintain browser storage like a real page.
- No automatic resource loading: script and image tags do not cause their resources to be fetched.
When browser execution is needed, Cheerio’s introduction points to Puppeteer or Playwright; it also names jsdom for DOM emulation. A useful workflow is to render with a browser tool, then pass the resulting HTML to Cheerio for extraction.
Complete runnable scraper
Save as extract.mjs, then run node extract.mjs https://example.com/.
import * as cheerio from 'cheerio';
const target = process.argv[2];
if (!target) throw new Error('Usage: node extract.mjs <url>');
const response = await fetch(target, {
headers: { 'user-agent': 'cheerio-example/1.0' },
signal: AbortSignal.timeout(15000)
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const contentType = response.headers.get('content-type') || '';
if (!/html|xml/i.test(contentType)) throw new Error(`Expected HTML or XML, got ${contentType || 'unknown content type'}`);
const $ = cheerio.load(await response.text());
const result = {
title: $('title').first().text().trim(),
headings: $('h1, h2, h3').map((_, el) => $(el).text().trim()).get(),
links: $('a[href]').map((_, el) => ({ text: $(el).text().trim(), href: $(el).attr('href') })).get()
};
console.log(JSON.stringify(result, null, 2));
Cheerio versus a browser tool
| Requirement | Cheerio | Browser automation |
|---|---|---|
| Parse response HTML | Yes | Yes |
| CSS selection and traversal | Yes | Yes |
| Run page JavaScript | No | Yes |
| Layout, screenshots, canvas, computed styles | No | Yes |
| Lightweight server-side transformation | Yes | Usually excessive |
Use Cheerio when markup is available and the job is extraction or transformation. Use Puppeteer or Playwright when the page must execute scripts or behave like a browser. Use jsdom when DOM emulation is the requirement.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Selector is empty | Content is client-injected, selector is wrong, or markup differs. | Inspect the raw response, verify selector, or render with Puppeteer/Playwright. |
| Text is missing | Value may be in an attribute or form property. | Read attr('content'), attr('value'), or the correct element. |
fromURL rejects response |
Content-Type is not HTML/XML, such as PDF or JSON. | Check headers and use a format-specific parser or validate a manual fetch. |
| Unexpected nesting | parse5 applies HTML tree-construction rules. | Consider XML/htmlparser2 mode where suitable and adjust selectors to the parsed tree. |
| Garbled characters | Bytes were decoded with the wrong encoding. | Use loadBuffer or decodeStream for encoding sniffing. |
| CAPTCHA or login page | Server returned an anti-bot or authenticated response. | Respect access controls and use an approved authentication/browser flow. |
Performance, reliability, and cost
- Performance: parse once, keep selectors specific, extract only needed fields, and reuse the loaded object.
- Memory: avoid retaining full documents across large batches; emit compact records and release references. Use htmlparser2 when its documented trade-offs fit.
- Reliability: set network timeouts, check status and content type, handle redirects/retries deliberately, and record each source URL.
- Cost: Cheerio is an npm dependency. Operational costs are network, compute, storage, and any browser service used before parsing; Cheerio provides no hosted scraping quota.
Or skip the browser setup
If you need a clean screenshot rather than HTML extraction, ScreenshotNeo provides a single API request. It accepts cookie and consent banners, removes 60+ known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
It supports full-page or element captures, dark mode, device presets, custom viewports and retina scale, PDF output, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs, bulk capture, usage reporting, and an MCP server for AI agents. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The MCP server gives AI agents screenshot tools. Free usage is 1,000 screenshots monthly with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Is Cheerio a browser?
No. It parses supplied markup and exposes a jQuery-like API; it does not render or run JavaScript.
Can Cheerio scrape a React site?
Only the content in the response. Render first with a browser tool when the data appears after client execution.
Does Cheerio support XML?
Yes. XML mode uses htmlparser2 behavior.
Can I use Cheerio in the browser?
Cheerio is chiefly used in Node.js for server-side parsing. For browser DOM work, native DOM APIs or a browser-focused library are generally more suitable.


