ScreenshotNeo

BlogGuides

The Ultimate Puppeteer Web Scraping Guide for 2026

Scrape JavaScript-rendered pages with Puppeteer using reliable waits, selectors, request handling, and practical guidance on reliability and compliance.

By the ScreenshotNeo team29 September 202610 min read

The Ultimate Puppeteer Web Scraping Guide for 2026

Puppeteer scrapes a JavaScript website by opening it in a real browser, waiting for the page state your extraction depends on, and reading the rendered DOM. Install puppeteer, build a bounded browser job, prefer Locators and state-based waits over fixed sleeps, validate every extracted record, and close browser resources in finally. Use Puppeteer when browser execution is necessary; for stable HTML or an authorized API, a plain HTTP client is usually simpler.

This guide covers Puppeteer setup, a runnable scraper, selector and wait strategies, pagination, API observation, request interception, production reliability, troubleshooting, and the legal and ethical checks that belong in a scraping workflow.

1. What Puppeteer does and when to use it

Puppeteer is a JavaScript library that automates Chrome and Firefox through the Chrome DevTools Protocol and WebDriver BiDi. Its documented uses include navigation, screenshots, PDF generation, UI testing, and performance analysis. Scraping is another application: browser execution allows a scraper to read content produced after JavaScript runs. Chrome for Developers: Puppeteer

Choose the lightest access method that provides the authorized data:

  • HTTP client: use when the needed content is in stable HTML or an authorized endpoint. It starts quickly and avoids browser rendering.
  • Puppeteer: use when a page requires JavaScript, browser state, or interaction to expose the content you are allowed to collect.
  • Do not use browser automation to defeat controls: CAPTCHAs, paywalls, authentication boundaries, and explicit technical restrictions require an authorized route.

2. Install Puppeteer and prepare the browser

The puppeteer package downloads a compatible Chrome during installation. puppeteer-core does not manage a browser download, so select and manage the browser yourself. Pin a major version in the lockfile and record the browser revision in deployment metadata so upgrades are reviewable. If installation scripts are blocked by your package manager, allow the install script or run npx puppeteer browsers install. Puppeteer installation guide

Puppeteer executes the page before the scraper reads and validates rendered content.
Puppeteer executes the page before the scraper reads and validates rendered content.
npm init -y
npm install puppeteer

Save the following as scrape.mjs and run it with Node.js. Set TARGET_URL to a site and path you are permitted to access. The example deliberately uses semantic attributes and validates missing fields instead of silently producing shifted or incomplete rows.

import puppeteer from 'puppeteer';

const targetUrl = process.env.TARGET_URL ?? 'https://example.com/';
const navigationTimeoutMs = 30_000;
const jobTimeoutMs = 60_000;

const browser = await puppeteer.launch({ headless: true });
let page;
let timer;

try {
  page = await browser.newPage();
  await page.setViewport({ width: 1365, height: 900 });
  await page.setDefaultNavigationTimeout(navigationTimeoutMs);
  await page.setDefaultTimeout(10_000);

  // Bound the whole job as well as individual browser operations.
  const deadline = new Promise((_, reject) => {
    timer = setTimeout(() => reject(new Error('Overall job deadline exceeded')), jobTimeoutMs);
  });

  const scrape = async () => {
    const response = await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
    if (!response) throw new Error('Navigation returned no HTTP response');
    if (response.status() >= 400) throw new Error(`HTTP ${response.status()}`);

    // Replace this with a selector that marks the page's actual ready state.
    await page.locator('body').wait();

    const result = await page.evaluate(() => {
      const title = document.querySelector('h1')?.textContent?.trim() ?? null;
      const description = document.querySelector('main p')?.textContent?.trim() ?? null;
      return {
        title,
        description,
        finalUrl: location.href,
        retrievedAt: new Date().toISOString(),
      };
    });

    if (!result.title) throw new Error('Required field missing: h1 title');
    return result;
  };

  const record = await Promise.race([scrape(), deadline]);
  console.log(JSON.stringify(record, null, 2));
} finally {
  clearTimeout(timer);
  if (page) await page.close().catch(() => {});
  await browser.close();
}

The example is a starting adapter, not a universal selector recipe. Replace h1 and main p with selectors tied to the target’s stable structure. In a worker process, launch one browser and create isolated BrowserContext instances for jobs that need separate cookies or storage. Keep each target’s URL construction, selectors, pagination, normalization, and validation in a site-specific adapter.

3. Selectors and waits that resist flakiness

Puppeteer recommends Locators for interactions. A Locator waits for an element to exist and reach the appropriate state, and Puppeteer supports CSS, XPath, text, accessibility, and Shadow DOM selector syntax. Prefer selectors based on stable IDs, semantic roles, labels, or documented data attributes. Avoid selectors that depend on generated class names or a long chain of layout containers. Puppeteer page interactions

Need Useful mechanism Watch for
Element appears page.locator(selector).wait() or page.waitForSelector() Presence is not always visibility or completed content.
Application condition becomes true page.waitForFunction(predicate) Keep the predicate specific and the timeout bounded.
Known request or response arrives page.waitForRequest() or page.waitForResponse() Register the wait before triggering the action.
Network becomes quiet page.waitForNetworkIdle() Long polling, analytics, or streaming can prevent idleness.
Fixed delay page.waitForTimeout() Use only when a known timer is part of the page behavior; delays are either wasteful or too short.

domcontentloaded means the initial document has been parsed; it does not mean a client-rendered application has finished fetching and drawing its data. load waits for load-event resources and can be unnecessarily slow. Network idle can be useful for finite applications, but it is not a universal readiness signal. Choose the condition that corresponds to the data you will extract.

For a click that causes navigation, start waiting before clicking to avoid a race:

const [response] = await Promise.all([
  page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
  page.locator('a.next').click(),
]);
// response may be null for History API or same-document navigation.

The navigation API notes that same-document and History API navigations can resolve with null. Puppeteer waitForNavigation API

4. Extract records, paginate, and preserve provenance

Run DOM reads in page.evaluate() and return plain data. Normalize and validate in Node.js so your scraper can distinguish an absent field from a valid empty string. Resolve links against the page URL, preserve the source URL and retrieval time, and parse locale-sensitive dates or prices deliberately. For embedded JSON, select the expected script element and handle parse errors; do not assume every script block has the same schema.

const records = await page.evaluate(() => {
  return [...document.querySelectorAll('[data-product-id]')].map((card) => {
    const link = card.querySelector('a[href]');
    return {
      id: card.getAttribute('data-product-id'),
      name: card.querySelector('[data-name]')?.textContent?.trim() ?? null,
      href: link ? new URL(link.href, location.href).href : null,
    };
  });
});

for (const row of records) {
  if (!row.id || !row.name || !row.href) {
    throw new Error(`Invalid product record: ${JSON.stringify(row)}`);
  }
}

For pagination, identify whether the control changes the URL, updates history, or replaces content in place. Wait for the corresponding URL, response, or content marker; then deduplicate by a stable record key and stop on an explicit end condition. Set a maximum page count as a safety bound, and do not blindly retry a form submission or any action with side effects.

5. Observe API responses and control network traffic

If the page loads data through an API, observe the request or response that the page already makes, and parse it only if you are authorized to access that data. Match narrowly on URL and response shape; sites may change their internal endpoints without notice. Respect authentication, published rate limits, and access rules.

const responsePromise = page.waitForResponse((response) =>
  response.url().includes('/api/catalog') && response.request().method() === 'GET'
);
await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
const apiResponse = await responsePromise;
if (!apiResponse.ok()) throw new Error(`Catalog API returned ${apiResponse.status()}`);
const payload = await apiResponse.json();

Request interception can reduce transferred assets when images, fonts, analytics, or known third-party resources are not needed. Start by preserving documents, scripts, stylesheets, XHR/fetch, and resources essential to the target application. Measure whether the page still renders correctly before expanding a block list. Once interception is enabled, every intercepted request must be continued, responded to, aborted, or otherwise resolved; a stalled request can stall the page. Puppeteer network interception guide

await page.setRequestInterception(true);
page.on('request', (request) => {
  const type = request.resourceType();
  const block = ['image', 'font'].includes(type);
  if (block) return request.abort();
  return request.continue();
});

Attach the handler before navigation. If another handler or library also manages interception, coordinate ownership and check the request’s handled state according to the installed Puppeteer version. Do not use interception to evade a site’s security or access controls.

6. Production architecture and operational reliability

  1. Pin and record versions. Keep the Puppeteer package and browser revision reproducible in deployment. The current getting-started documentation is labeled Puppeteer 25.12.0; check the docs matching the version you install. Puppeteer getting started
  2. Reuse the browser process. Launch one browser per worker process, then create and close pages or isolated contexts per job. This avoids repeatedly paying browser startup overhead while preserving cookie separation where needed.
  3. Bound work. Set navigation, selector, and overall job deadlines. Record status, final URL, elapsed time, and a compact error category. Avoid infinite retries.
  4. Retry carefully. Retry idempotent page loads for transient network failures with exponential backoff and jitter. Do not automatically replay submissions or other state-changing actions.
  5. Limit concurrency. Keep request rates within the target site’s tolerated rate. Add queues and per-host limits instead of opening uncontrolled pages.
  6. Manage memory. Close pages and contexts in finally; recycle workers or pages when their memory use grows. Save HTML or response payloads only when permitted and necessary, and redact personal data.
  7. Detect wrong-but-successful pages. A 200 response may still be a consent page, expired login, soft 404, empty result, or bot check. Validate page identity and required fields, not just HTTP status.

7. Troubleshooting common Puppeteer scraping errors

Symptom Likely cause Fix
Selector timeout Wrong selector, wrong page, or content not ready Check final URL and a screenshot/DOM artifact; wait for the actual content marker and validate the selector in the current page.
Navigation timeout Slow target, blocked request, or waiting for an event that never fires Use a bounded timeout and a more relevant state such as domcontentloaded; inspect status and failed requests.
Click sometimes misses the next page Navigation wait registered after the click Use Promise.all with waitForNavigation started first; handle same-document updates separately.
Network idle never arrives Long polling, analytics, streaming, or persistent requests Wait for a specific selector, response, or application predicate instead.
Request hangs after enabling interception A request handler did not resolve every intercepted request Ensure every branch continues, aborts, or responds; log request URL and resource type during diagnosis.
Chrome executable missing Install scripts were disabled or puppeteer-core is used without a browser path Allow the Puppeteer install script or run npx puppeteer browsers install; for core, configure a compatible browser explicitly.
Browser crashes under load Too many concurrent pages, large pages, or leaked contexts Reduce concurrency, close resources in finally, and recycle workers on a controlled schedule.
Empty data despite a successful load Soft 404, changed markup, login expiry, consent overlay, or API schema change Validate page identity and required fields; capture diagnostics and update the site adapter only after inspecting the cause.

8. Performance, cost, and alternatives

Browser scraping costs more CPU and memory than fetching HTML because it starts or reuses a browser and executes page resources. Reuse browser processes, block only resources that are unnecessary and safe to omit, keep concurrency bounded, and avoid waiting for events unrelated to the extraction. Cache immutable results only where site terms and freshness requirements allow. There is no universal throughput figure: page weight, JavaScript behavior, browser environment, and target limits determine performance.

A screenshot API can prepare a clean capture without requiring a local browser workflow.
A screenshot API can prepare a clean capture without requiring a local browser workflow.

For screenshot work, the simplest option may be to request an image rather than operate a browser. ScreenshotNeo is a website screenshot API and MCP server: one GET request returns a PNG, JPEG, WebP, or PDF. Its screenshot controls include full-page capture, element selection, viewport presets, custom CSS and JavaScript, waits, and caching; consult the ScreenshotNeo API documentation for parameters and response details. The request below captures a page to WebP:

Or skip the browser setup

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

There is no single yes-or-no answer for every target, dataset, and jurisdiction. Check the site’s terms, copyright and database rights, privacy requirements, authentication boundaries, contractual limits, and rate expectations before collection. RFC 9309 standardizes the robots.txt protocol: when successfully fetched, parseable rules should be followed, but the RFC explicitly says those rules are not access authorization. Robots rules are one input to the review, not permission to access protected material. RFC 9309

When personal data is involved, document the purpose and legal basis, minimize collection, set retention limits, and get appropriate legal review. The European Data Protection Board’s 2026 web-scraping materials discuss GDPR legal bases and special-category data. Puppeteer’s own security policy also places responsibility on the calling code to use browser installation, automation, and inspection safely and as intended. EDPB Guidelines 03/2026 consultation · Puppeteer security policy

10. Frequently asked questions

Can Puppeteer scrape sites that require login?

Only when you have authorization to access and collect the relevant content. Use a dedicated account and isolated browser context where appropriate, protect credentials, and do not cross authentication or access boundaries.

Should I use Puppeteer or an official API?

Use an official, authorized API when it exposes the required data and terms permit your use. It generally avoids rendering overhead and DOM fragility. Choose Puppeteer when browser execution is necessary.

Can Puppeteer scrape PDFs?

Puppeteer can generate PDFs from pages, but scraping text from an existing PDF is a separate document-extraction task. Choose a parser appropriate to the PDF and verify the publisher permits the intended use.

How do I keep selectors working after a redesign?

Prefer semantic roles, labels, stable IDs, or documented data attributes. Put selectors in a site adapter, validate required fields, and treat missing data as an observable failure that calls for review.

Can I run several scraper jobs at once?

Yes, with bounded concurrency. Isolate cookies and storage when jobs require it, limit work per host, and monitor memory and failure rates before increasing parallelism.