Web Scraping and Browser Automation with Crawlee
Learn when to use CheerioCrawler or PlaywrightCrawler, install Crawlee correctly, handle sessions and proxies, and build reliable scraping workflows.

Direct answer: use Crawlee’s CheerioCrawler when the data is already present in the HTML returned by an HTTP request. Use PlaywrightCrawler when the page needs JavaScript execution, browser events, scrolling, clicks, or other browser behavior. Use PuppeteerCrawler when your team already uses Puppeteer or prefers its API. Crawlee supplies the crawling, queue, retry, session, proxy, and storage building blocks; browser automation libraries are separate dependencies.
Crawlee is an open-source web scraping and browser automation library with JavaScript and Python implementations. This guide focuses on the JavaScript package. The current JavaScript documentation is version 3.18, and the official changelog lists 3.18.1 on August 12, 2026 and 3.18.0 on August 4, 2026. Check the live changelog before pinning a version.
What is Crawlee?
Crawlee provides a common way to define requests, process responses or browser pages, retry failures, limit a crawl, maintain sessions, route requests, and write results to datasets. It can run on your own machine or on other cloud infrastructure. Apify is an optional managed deployment path, not a requirement. The project is open source under the Apache License 2.0, as documented in the repository README.
The important architectural choice is the crawler class. An HTTP crawler downloads a response and parses it. A browser crawler launches a browser, loads the page, runs JavaScript, and exposes browser controls. Starting with the lightest crawler that can produce correct data usually reduces installation size, memory use, and operational complexity.
CheerioCrawler vs PlaywrightCrawler vs PuppeteerCrawler
| Requirement | Start with | What it does | Constraint |
|---|---|---|---|
| Server-rendered or static HTML | CheerioCrawler |
HTTP requests plus Cheerio HTML parsing | It cannot render client-side JavaScript |
| JavaScript application, interaction, or browser APIs | PlaywrightCrawler |
Controls a browser through Playwright | Requires the Playwright dependency and browser binaries |
| Existing Puppeteer code | PuppeteerCrawler |
Controls a browser through Puppeteer | Requires Puppeteer and its browser setup |
Do not choose a browser crawler merely because it sounds more powerful. Browser execution is appropriate when it is required for correctness. For a page whose product cards are in the initial HTML, CheerioCrawler is simpler. For a page that fills those cards after an API call, uses an infinite-scroll list, or requires a click to reveal content, use PlaywrightCrawler or PuppeteerCrawler.

Install Crawlee and the browser dependency
The JavaScript quick start requires Node.js 16 or later. Create a project and install the base package:
mkdir crawlee-example
cd crawlee-example
npm init -y
npm install crawlee
Browser crawlers need an explicit browser automation install. Crawlee does not bundle Playwright or Puppeteer:
# Choose one browser automation library
npm install crawlee playwright
# or
npm install crawlee puppeteer
If you want smaller, focused packages, the API documentation also lists @crawlee/cheerio and @crawlee/playwright. The beginner-friendly project generator is:
npx crawlee create my-crawler
Select a starter template, then set a request limit while developing so a selector mistake cannot create an unbounded crawl.
How do I scrape a website with Crawlee?
1. Scrape server-rendered HTML with CheerioCrawler
This complete example requests two pages, extracts titles and links, limits the crawl, and stores results in Crawlee’s default dataset.
import { CheerioCrawler } from 'crawlee';
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 20,
async requestHandler({ request, $, enqueueLinks, pushData, log }) {
const title = $('title').first().text().trim();
const links = $('a[href]')
.map((_, element) => $(element).attr('href'))
.get()
.filter(Boolean);
await pushData({
url: request.url,
title,
linkCount: links.length,
});
log.info(`Saved ${request.url}`);
await enqueueLinks({ selector: 'a[href]' });
},
});
await crawler.run([
'https://example.com/',
]);
Run it as an ES module by adding "type": "module" to package.json, or place the code in a module-compatible file. The handler receives a Cheerio selector function, not a browser page. There is no JavaScript execution, DOM layout, click, or scroll.
2. Render JavaScript with PlaywrightCrawler
When content appears only after the page runs JavaScript, use a browser crawler. Install Playwright separately, then use the page object in the handler:
import { PlaywrightCrawler } from 'crawlee';
const crawler = new PlaywrightCrawler({
maxRequestsPerCrawl: 10,
async requestHandler({ request, page, enqueueLinks, pushData, log }) {
await page.waitForLoadState('domcontentloaded');
// Use a page-specific readiness condition when possible.
await page.locator('[data-product-card]').first().waitFor({
state: 'visible',
timeout: 15000,
}).catch(() => log.warning('Product selector was not found'));
const title = await page.title();
const products = await page.locator('[data-product-card]').evaluateAll(
cards => cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null,
})),
);
await pushData({ url: request.url, title, products });
await enqueueLinks({ selector: 'a[href]' });
},
});
await crawler.run(['https://example.com/catalog']);
Replace the example selectors with selectors from the target site. A wait for a meaningful element is generally more reliable than a fixed delay because it expresses the condition your extraction needs.
3. Use PuppeteerCrawler when Puppeteer is your standard
import { PuppeteerCrawler } from 'crawlee';
const crawler = new PuppeteerCrawler({
maxRequestsPerCrawl: 10,
async requestHandler({ request, page, pushData }) {
await page.waitForSelector('main');
await pushData({
url: request.url,
title: await page.title(),
text: await page.$eval('main', element => element.textContent?.trim() ?? ''),
});
},
});
await crawler.run(['https://example.com/']);
Configuration that matters in production
Request limits and queues
Set maxRequestsPerCrawl during development and for deliberately bounded jobs. For larger jobs, enqueue URLs from a sitemap, seed list, or discovered links and make the queue the source of truth. Keep URL normalization and deduplication consistent so query-string variants do not create accidental duplicates.
Waiting for the right state
Browser pages can be technically loaded while their useful data is still pending. Prefer, in order:
- Wait for a selector that proves the data is present.
- Wait for a known application state or response when the site exposes one.
- Use network-idle or a short delay only when the page has no better readiness signal.
Always put an upper timeout on waits. A missing selector should become a classified failure or an empty result according to your data contract, rather than holding a browser indefinitely.
Retries and idempotency
Retries help with transient network and browser failures. Make handlers idempotent: writing the same item twice should either overwrite the same key or be detectable downstream. Store the source URL and a stable record identifier with every item. Do not assume a retry means the target page changed; it may simply be a temporary connection or rendering failure.
Proxies and sessions
Crawlee’s proxy configuration can be integrated with HTTP and browser crawler classes. Its SessionPool associates sessions with cookies, proxy details, and other session-specific settings. These mechanisms help you model separate sessions and retain state; they do not guarantee anonymity, access, successful challenge completion, or permission to collect data. Follow the target site’s terms, access controls, and applicable law.
A session is useful when a site keeps a login or preference in cookies. A proxy configuration is useful when your deployment has an approved pool of network egress locations. Keep credentials outside source code, rotate them according to your provider’s rules, and record which session produced each item so failures can be investigated.
Headers, authentication, and browser context
For HTTP crawling, set request headers only when you have a legitimate reason, such as an API’s documented authorization scheme. For browser crawling, configure the browser context for the required locale, timezone, or authenticated state. Treat cookies and authorization tokens as secrets and avoid logging them. A custom user agent should identify your client accurately rather than pretending to be an unrelated browser.
Storage and exports
pushData writes extracted records to a dataset. Keep raw evidence such as the URL, retrieval timestamp, and a compact failure reason alongside normalized fields. Export only the fields consumers need. For repeatable pipelines, define a schema and validate required fields before writing them.
Common edge cases
- Data is absent in Cheerio: inspect the raw response. If the HTML contains only an application shell, switch to PlaywrightCrawler or call the documented data endpoint directly.
- Infinite scroll: scroll in a bounded loop, wait for a count or end marker, and stop after a maximum number of pages. Do not rely on one arbitrary sleep.
- Cookie or consent dialogs: locate the dialog and click the permitted choice before extracting data. Keep this behavior site-specific and test it against localization variants.
- Changing selectors: prefer stable attributes or semantic structure. Store a sample of failed pages so selector changes are visible.
- Downloads and PDFs: decide whether the result belongs in browser automation or a direct HTTP download workflow. Do not parse a download as if it were an HTML document.
- Authentication expiry: detect redirects to login and refresh the approved session rather than saving the login page as valid data.
- Locale-dependent content: set locale and timezone deliberately, then record them with the result so values can be reproduced.
Performance, reliability, and cost planning
CheerioCrawler generally uses fewer resources because it performs HTTP and HTML parsing without launching a browser. Browser crawlers incur browser startup, page memory, JavaScript execution, and asset-loading costs. The official material describes CheerioCrawler as fast and efficient but does not provide a measured benchmark; do not plan capacity from an invented requests-per-second number.
Improve throughput by using Cheerio where it is correct, limiting browser work to routes that need it, blocking unnecessary resources when the page still renders correctly, and reusing a crawler process for a batch. Keep concurrency within the memory and CPU limits of your host. Monitor queue depth, request duration, retry counts, browser crashes, and records rejected by schema validation.
Reliability comes from bounded work and observable failure states. Set request and navigation timeouts, capture the URL and error class, retain enough response or screenshot evidence to debug, and alert on changes in success rate or extracted-field completeness. A proxy or session pool can change routing and state, but it cannot correct a broken selector or make an inaccessible site available.
Your direct cost depends on where Crawlee runs: local compute, a cloud VM, containers, browser binaries, proxy services, storage, and any optional managed platform. Apify can provide managed deployment and scheduling, but Crawlee also runs locally or on other cloud infrastructure. Budget browser jobs separately from lightweight HTTP jobs.
Troubleshooting Crawlee
| Symptom | Likely cause | Fix |
|---|---|---|
Cannot find module 'playwright' |
Browser dependency was not installed | Run npm install crawlee playwright (or the Puppeteer equivalent) and install required browser binaries. |
| Selector returns zero elements | Wrong crawler type, changed markup, or page not ready | Inspect raw HTML; switch to a browser crawler if JavaScript fills the DOM; wait for a stable readiness selector. |
| Navigation timeout | Slow page, blocked resource, or an application that never reaches the chosen load state | Use a suitable timeout and readiness condition, block irrelevant resources, and record the failing URL for review. |
| Repeated retries never succeed | Persistent selector or authorization error rather than a transient failure | Classify the error, stop retrying that request, and inspect the saved response or browser state. |
| Duplicate records | Multiple URL forms or non-idempotent writes | Normalize URLs, deduplicate queue entries, and write a stable source key. |
| Memory grows during a browser crawl | Too much concurrency, unclosed pages, or loading unnecessary assets | Reduce concurrency, ensure handlers finish, limit crawl scope, and block resources that are irrelevant to extraction. |
| Session appears logged out | Expired cookies, a new session, or a login redirect | Detect the redirect, refresh the approved authentication state, and persist session diagnostics. |
Or skip the browser setup
If your goal is a clean image or PDF of a URL rather than structured extraction, ScreenshotNeo provides a single GET request. It accepts the consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for the complete option list. The service supports full-page capture with lazy images loaded, CSS element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper sizes and margins, landscape mode and page ranges, HTML/CSS to image, custom CSS and JavaScript, clicks, selector or network-idle waits, ad/tracker/request/resource blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', buffer));
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Plans include every feature: 1,000 shots per month free with no card, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Start with 1,000 free screenshots a month—no card required.
FAQ
Can Crawlee scrape pages that require JavaScript?
Yes. Use PlaywrightCrawler or PuppeteerCrawler and wait for the page state your extraction needs. CheerioCrawler cannot execute client-side JavaScript.

Do I need Apify to use Crawlee?
No. Crawlee runs locally and on other cloud infrastructure. Apify is an optional platform for managed deployment and related tooling.
Are Playwright and Puppeteer included with Crawlee?
No. Install the browser automation library separately, such as npm install crawlee playwright or npm install crawlee puppeteer.
Will a proxy guarantee that a site allows my crawler?
No. Proxy and session features configure routing and state; they do not guarantee access or override a site’s rules.
How should I choose between browser automation and a screenshot API?
Use Crawlee when you need structured data, links, business rules, and a crawler queue. Use a screenshot API when the deliverable is a rendered image or PDF and you want to avoid maintaining browser setup and cleanup logic.


