What Is Puppeteer in Web Scraping?
Puppeteer controls Chrome or Firefox from JavaScript, letting scrapers load, interact with, and extract content from rendered pages.
Puppeteer is a JavaScript library for controlling Chrome or Firefox programmatically. In web scraping, it opens a real browser, waits for JavaScript to run, interacts with page elements, and reads the rendered DOM. That makes it useful for single-page applications and other sites whose content is not present in the initial HTML.
Puppeteer is a browser-automation library, not a scraping service or prebuilt dataset. It is also used for UI testing, screenshots, PDF generation, tracing, performance analysis, and pre-rendering. The project describes it as a high-level API for controlling Chrome or Firefox over the Chrome DevTools Protocol (CDP) or WebDriver BiDi.
Official overview: What is Puppeteer?
How Puppeteer fits into a scraping workflow
- Your Node.js program launches a browser.
- Puppeteer creates a browser context and page.
- The page navigates to a URL and executes its JavaScript.
- Your code waits for a useful condition, such as a selector or network activity settling.
- Puppeteer queries the DOM, clicks controls, fills forms, or scrolls to trigger lazy loading.
- Your program normalizes and stores the extracted data, then closes the browser.
Because a real browser executes page code, Puppeteer can handle workflows that a simple HTTP client cannot. It does not grant permission to access a website, bypass access controls, defeat bot checks, or ignore robots policies and rate limits. Use it only against sites and data you are authorized to access; the calling code remains responsible for safe and intended use.
Install Puppeteer
The normal package downloads a compatible Chrome build during installation:
npm install puppeteer
If you already manage the browser binary yourself, install puppeteer-core. It provides the library without downloading Chrome:
npm install puppeteer-core
Package managers can block install scripts. If the browser was not downloaded, install a compatible browser manually with:
npx puppeteer browsers install chrome
Read the current installation guide before pinning versions in production because browser revisions and supported defaults change.
A complete Puppeteer scraping example
This script loads a page, waits for article cards, extracts text and links, writes JSON, and closes the browser even when an error occurs.
const puppeteer = require('puppeteer');
const fs = require('node:fs/promises');
async function scrape() {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
await page.goto('https://example.com/articles', {
waitUntil: 'domcontentloaded',
timeout: 30000
});
await page.waitForSelector('article.card', { timeout: 15000 });
const records = await page.$$eval('article.card', cards => cards.map(card => ({
title: card.querySelector('h2')?.textContent?.trim() ?? null,
url: card.querySelector('a')?.href ?? null,
summary: card.querySelector('.summary')?.textContent?.trim() ?? null
})));
await fs.writeFile('articles.json', JSON.stringify(records, null, 2));
console.log(`Saved ${records.length} records`);
} finally {
await browser.close();
}
}
scrape().catch(error => {
console.error(error);
process.exitCode = 1;
});
Replace the URL and selectors with selectors from the site you are authorized to process. Prefer stable attributes such as data-testid over classes that are primarily used for styling.
Wait for the right condition
waitUntil: 'domcontentloaded' only means the initial document was parsed. For rendered content, add a page-specific wait:
await page.waitForSelector('[data-testid="product-list"]');
For a page that finishes through background requests, you can wait for network idle, but long polling and analytics can prevent it from becoming idle:
await page.goto(url, { waitUntil: 'networkidle2', timeout: 45000 });
A selector wait is usually more deterministic than an arbitrary delay. Use a delay only when the site has a known animation or delayed update:
await new Promise(resolve => setTimeout(resolve, 1000));
Click, paginate, and scroll
await page.click('button.load-more');
await page.waitForSelector('.new-items');
await page.evaluate(async () => {
window.scrollTo(0, document.body.scrollHeight);
});
For infinite scroll, loop until the item count stops increasing or a documented end condition appears. Always set a maximum number of iterations.
Read content from inside the page
const result = await page.evaluate(() => ({
title: document.title,
description: document.querySelector('meta[name="description"]')?.content ?? null,
text: document.body.innerText
}));
page.evaluate runs in the page context. Keep data extraction separate from application logic and validate fields before writing them to a database.
Browser, context, and page options
| Option | Purpose | Practical note |
|---|---|---|
headless |
Run without a visible window. | Use headless mode on servers; set false while debugging locally. |
executablePath |
Use a browser installed by your OS or container. | Useful with puppeteer-core; keep the browser and library compatible. |
args |
Pass Chromium launch flags. | Only add flags required by your environment and understand their security impact. |
userAgent |
Set the page’s user-agent string. | Identify your automation honestly where required by the site. |
setExtraHTTPHeaders |
Add request headers. | Do not place secrets in logs or expose them to untrusted pages. |
| Cookies and storage | Reuse an authenticated session. | Store session data securely and minimize its lifetime. |
| Request interception | Inspect, block, or rewrite requests. | Blocking fonts, images, ads, and analytics can reduce work but may change page behavior. |
await page.setExtraHTTPHeaders({ 'Accept-Language': 'en-US,en;q=0.9' });
await page.setCookie({
name: 'consent',
value: 'accepted',
domain: 'example.com',
path: '/'
});
Scraping single-page applications
Client-rendered applications often return a small HTML shell and populate the page after JavaScript runs. Puppeteer can wait for the rendered component:
await page.goto('https://example.com/app', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('#app [data-loaded="true"]', { timeout: 30000 });
const rows = await page.$$eval('#app .row', nodes => nodes.map(node => node.textContent.trim()));
When the application calls a documented JSON endpoint, calling that endpoint directly is usually simpler and cheaper than rendering a browser. Use Puppeteer when browser execution or interaction is part of the required behavior.
HTTP clients versus Puppeteer
A direct HTTP request can be sufficient for server-rendered HTML. These examples show the distinction; they do not use Puppeteer.
cURL
curl --fail --location --max-time 30 https://example.com/articles
Python
import requests
response = requests.get('https://example.com/articles', timeout=30)
response.raise_for_status()
print(response.text)
Neither request executes the page’s JavaScript. If the data appears only after rendering, use a browser automation library such as Puppeteer or an authorized data endpoint.
Puppeteer and browser support
From Puppeteer v23 onward, the project supports Chrome and Firefox. Chrome uses CDP by default; Firefox uses WebDriver BiDi by default. Puppeteer continues to support CDP for Chrome, including Chrome-specific capabilities and existing automation. Check the current FAQ and release notes before deploying a particular browser revision.
Puppeteer versus Selenium
| Question | Puppeteer | Selenium |
|---|---|---|
| Primary fit | Node.js browser automation through CDP and WebDriver BiDi. | Browser automation with language bindings across a broader ecosystem. |
| Languages | JavaScript/TypeScript is the core experience. | Useful when a team needs several language bindings. |
| Large-scale orchestration | Compose your own workers and infrastructure. | Selenium Grid and related orchestration can matter for distributed fleets. |
| Browser protocols | CDP for Chrome and WebDriver BiDi defaults described above. | Uses WebDriver-based interfaces and ecosystem tooling. |
Choose based on language requirements, browser and protocol needs, and whether you need centralized orchestration at scale. There is no universal winner.
Reliability and performance checklist
- Set navigation and selector timeouts explicitly.
- Use one browser process with multiple isolated contexts when that matches your workload.
- Close pages, contexts, and browsers in
finallyblocks. - Retry transient navigation failures with bounded exponential backoff.
- Record URL, status, elapsed time, selector failures, and browser errors.
- Limit concurrency to what the machine and the target site can handle.
- Block unnecessary resources only after confirming they are not required for rendering.
- Cache results when freshness requirements allow it.
- Respect the target site’s terms, access rules, rate limits, and privacy requirements.
Browser work consumes substantially more memory and CPU than parsing an HTTP response. Measure your own pages and workload instead of assuming a fixed throughput. Browser version, page complexity, media, network conditions, and concurrency all affect runtime.
Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
Could not find Chrome |
Install scripts were blocked or puppeteer-core was installed without a browser. |
Run npx puppeteer browsers install chrome, or set a valid executablePath. |
TimeoutError: Navigation timeout |
The page or network did not finish within the limit. | Raise the timeout carefully, use a narrower wait condition, and investigate slow resources. |
Waiting for selector failed |
The selector is wrong, content is gated, or rendering failed. | Inspect the page in headed mode, verify the selector, and wait for the actual application state. |
| Empty text or missing rows | Extraction ran before client rendering completed. | Wait for a stable selector or application signal before evaluating the DOM. |
| Browser crashes under load | Too many concurrent pages, memory-heavy pages, or leaked resources. | Reduce concurrency, close pages, recycle workers, and monitor memory. |
| Works locally but not in a container | Missing shared libraries, sandbox restrictions, fonts, or browser binary. | Use a supported base image, install required dependencies, and follow the deployment guidance for the chosen browser. |
| Repeated blocks or access denials | The site has access controls, rate limits, or bot protection. | Do not attempt to bypass them. Confirm authorization, slow requests, and use an official API where available. |
Or skip the browser setup
If your goal is a reliable screenshot rather than custom DOM extraction, ScreenshotNeo is the first screenshot API to try: it produces clean shots, bills only clean shots, and its lowest paid plan is $5.
One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and whether it was billed. It also provides an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently asked questions
Is Puppeteer only for scraping?
No. Its documented uses include testing, screenshots, PDFs, tracing, performance analysis, UI automation, and pre-rendering single-page applications.
Can Puppeteer scrape any website?
It can automate pages that your browser can access, but access rules, authentication, rate limits, terms, and privacy obligations still apply. Puppeteer does not make restricted collection permissible.
Should I use Puppeteer or a direct API?
Use a documented API or direct HTTP request when it supplies the data you need. Use Puppeteer when browser execution, interaction, or rendered state is required.
Does Puppeteer include a browser?
puppeteer normally downloads a compatible Chrome during installation. puppeteer-core does not, so you provide the browser yourself.
Why does a scraper return an empty page?
Most often the script reads the DOM before the application finishes rendering. Wait for a page-specific selector or state, then extract the content.


