ScreenshotNeo

BlogHow-to

How to Increase Web Scraping Speed with Puppeteer

Speed up Puppeteer scraping by removing unnecessary waits, reusing browsers, filtering requests, and tuning concurrency without sacrificing correctness.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: Puppeteer scrapers become faster when each page does less work and waits only for the signal that proves its data is ready. Reuse one browser process, run a bounded number of pages or browser contexts, abort assets your extractor does not need, replace fixed sleeps with selector/request/response waits, and measure CPU, memory, network, errors, and target-site limits before increasing concurrency.

Puppeteer controls Chrome or Firefox through the DevTools Protocol or WebDriver BiDi and runs headless by default. The official documentation is the reference for its Page, Browser, and deployment APIs: Puppeteer documentation.

1. Measure the current bottleneck first

Do not start by adding workers. Split one scrape into measurable phases:

  • Browser launch and connection
  • Page or context creation
  • Navigation and readiness wait
  • Data extraction
  • Optional scrolling, clicking, or pagination
  • Teardown and persistence

Record median and tail duration (for example, p50 and p95), timeout rate, bytes transferred, successful records per minute, CPU, memory, and open connections. Test with a representative URL set. A fast page can hide a slow JavaScript-heavy page, a long poll, or an anti-bot challenge.

import puppeteer from 'puppeteer';

const urls = ['https://example.com', 'https://example.org'];
const browser = await puppeteer.launch({headless: true});

for (const url of urls) {
  const started = performance.now();
  const page = await browser.newPage();
  try {
    await page.goto(url, {waitUntil: 'domcontentloaded', timeout: 30000});
    const title = await page.title();
    console.log(JSON.stringify({
      url,
      title,
      milliseconds: Math.round(performance.now() - started)
    }));
  } finally {
    await page.close();
  }
}
await browser.close();

2. Choose the earliest reliable readiness condition

Every wait should answer a specific question: “What proves the data I need is available?” Use the earliest reliable condition.

Condition Use it when Speed and correctness notes
domcontentloaded The required values are in the initial document. Usually faster than waiting for every asset.
A selector Client code inserts the record or table you extract. Wait for the exact element, ideally with a meaningful state or count.
A request The page fetches a known API request that contains your data. Often avoids waiting for unrelated rendering.
A response You need a particular HTTP response and its status or body. Validate status and payload before extracting.
Network idle Network quiescence is genuinely the readiness signal. Analytics, ads, polling, and long-lived connections can delay it.
await page.goto(url, {
  waitUntil: 'domcontentloaded',
  timeout: 30000
});
await page.waitForSelector('[data-product-price]', {timeout: 10000});

const price = await page.$eval(
  '[data-product-price]',
  element => element.textContent.trim()
);

Use a response wait when the page’s data comes from a predictable endpoint:

const responsePromise = page.waitForResponse(response =>
  response.url().includes('/api/products') && response.request().method() === 'GET'
);
await page.goto(url, {waitUntil: 'domcontentloaded'});
const response = await responsePromise;
if (!response.ok()) throw new Error(`Product API returned ${response.status()}`);
const data = await response.json();

3. Remove fixed sleeps

A fixed delay makes every URL pay the worst-case wait and can still race the page. Replace waitForTimeout-style pauses with a selector, request, response, navigation, or a short polling condition tied to the actual state you need.

// Fragile and slow:
await new Promise(resolve => setTimeout(resolve, 5000));

// Faster and deterministic:
await page.waitForFunction(
  () => document.querySelectorAll('.result').length >= 20,
  {timeout: 10000}
);

Keep a timeout on every event-specific wait. A timeout exposes a page that did not reach the expected state instead of silently returning incomplete records.

4. Block only requests your extraction does not need

Images, fonts, media, trackers, and advertising requests can consume bandwidth and rendering time. Intercept and abort them when the target site still produces the data you need without them. Blocking stylesheets can change layout-dependent selectors; blocking scripts can remove the code that produces your data. Validate request policies per site.

await page.setRequestInterception(true);
page.on('request', request => {
  const type = request.resourceType();
  if (['image', 'font', 'media'].includes(type)) {
    return request.abort();
  }
  return request.continue();
});

await page.goto(url, {waitUntil: 'domcontentloaded'});

Prefer resource-type rules over broad URL rules at first. Log aborted URLs while developing, then add narrowly scoped host or path rules only when you know they are safe.

5. Reuse the browser and isolate jobs

Launching Chrome for every URL adds process startup and memory overhead. Launch one browser per worker process, then create pages or browser contexts for individual jobs. A context gives a job isolated cookies, storage, and permissions while sharing the browser process.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({headless: true});

async function scrape(url) {
  const context = await browser.createBrowserContext();
  const page = await context.newPage();
  try {
    await page.goto(url, {waitUntil: 'domcontentloaded', timeout: 30000});
    return await page.$eval('main', element => element.innerText);
  } finally {
    await context.close();
  }
}

try {
  for (const url of urls) console.log(await scrape(url));
} finally {
  await browser.close();
}

Use a page per job when shared cookies and storage are intentional. Use contexts when jobs must not see one another’s state. Close pages and contexts in finally blocks so failed jobs do not leak targets.

6. Bound concurrency with a queue

There is no universal “correct” number of Puppeteer pages. Increase workers until CPU, memory, network bandwidth, error rate, browser stability, or the target site’s permitted rate becomes the bottleneck, then back off. Unbounded tabs usually reduce throughput by causing contention and timeouts.

import puppeteer from 'puppeteer';

const urls = /* your URL list */;
const workerCount = 4;
const browser = await puppeteer.launch({headless: true});
let next = 0;

async function worker() {
  while (true) {
    const index = next++;
    if (index >= urls.length) return;
    const page = await browser.newPage();
    try {
      await page.goto(urls[index], {
        waitUntil: 'domcontentloaded',
        timeout: 30000
      });
      const record = await page.$eval('main', el => ({
        text: el.innerText
      }));
      console.log(JSON.stringify({url: urls[index], record}));
    } catch (error) {
      console.error(JSON.stringify({url: urls[index], error: String(error)}));
    } finally {
      await page.close();
    }
  }
}

try {
  await Promise.all(Array.from({length: workerCount}, worker));
} finally {
  await browser.close();
}

For production, replace the simple counter with a queue that supports backpressure, retries, cancellation, and graceful shutdown. Keep separate limits for browser workers and requests to any one host.

7. Keep browser configuration consistent

Run headless explicitly in deployment configuration and use one known Chrome or Firefox executable. Centralize launch arguments and defaults so workers do not accidentally perform different setup work. An explicit executable path can also prevent a worker from downloading or resolving a different browser at runtime.

const browser = await puppeteer.launch({
  headless: true,
  executablePath: process.env.CHROME_PATH || undefined,
  args: ['--disable-dev-shm-usage']
});

Only add launch flags you understand and have validated in your environment. Browser flags can affect security, rendering, or compatibility.

8. Handle scrolling, lazy loading, and pagination efficiently

Full-page content may not exist until the page scrolls. Scroll in measured steps, stop when the document height stops growing, and wait for a selector or response after each action instead of sleeping for a fixed interval.

let previousHeight = 0;
for (let i = 0; i < 20; i++) {
  const height = await page.evaluate(() => document.body.scrollHeight);
  if (height === previousHeight) break;
  previousHeight = height;
  await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
  await page.waitForFunction(
    oldHeight => document.body.scrollHeight > oldHeight,
    {timeout: 5000},
    previousHeight
  ).catch(() => {});
}

If pagination is backed by an API, capture and request the API pages directly when permitted and when the response contains the same records. This can avoid repeated layout and paint work, but preserve authentication, rate limits, and the site’s terms.

9. Avoid deployment bottlenecks

Profile launch, page creation, navigation, extraction, and teardown separately. In Google Cloud Run, Puppeteer’s troubleshooting guide documents a case where CPU is disabled after an HTTP response; background browser work then appears to take minutes. Keep CPU allocated for the lifetime of background browser work in that deployment pattern.

Also check container memory, shared memory limits, DNS latency, outbound bandwidth, and connection limits. A scraper that is fast locally can stall in a container when CPU throttles or when several Chromium processes compete for memory.

10. Retry carefully and preserve correctness

  • Retry transient navigation failures with exponential backoff and jitter.
  • Do not retry a deterministic selector timeout indefinitely; record the URL and diagnostic state.
  • Recreate a page or context after a crashed target.
  • Keep idempotent output writes so a retry cannot duplicate records.
  • Respect robots rules, terms of service, authentication requirements, privacy obligations, and explicit rate limits.
async function withRetry(task, attempts = 3) {
  let lastError;
  for (let attempt = 0; attempt < attempts; attempt++) {
    try {
      return await task();
    } catch (error) {
      lastError = error;
      if (attempt === attempts - 1) break;
      const delay = 250 * 2 ** attempt + Math.random() * 150;
      await new Promise(resolve => setTimeout(resolve, delay));
    }
  }
  throw lastError;
}

11. Common errors and fixes

Symptom Likely cause Fix
Navigation timeout of ... ms exceeded Slow origin, blocked resource, or a wait condition that never completes. Set a realistic timeout, use domcontentloaded or a specific selector, and capture diagnostics. Do not blindly increase every timeout.
Waiting for selector ... failed Wrong selector, consent wall, bot check, or data rendered only after an action. Inspect HTML and screenshots, handle the required state, and wait for the actual response or selector.
ERR_CONNECTION_RESET or intermittent 5xx Network instability, origin throttling, or too much concurrency. Reduce per-host concurrency, add bounded retries, and respect rate limits.
Scraper becomes slower as jobs increase CPU or memory contention, too many pages, or connection saturation. Lower worker count, reuse the browser, and compare p95 latency and error rate.
Data is missing after request blocking A blocked script, stylesheet, or API response was required. Log aborted requests and allow the required resource type or host.
Network-idle wait never finishes Analytics, polling, ads, WebSockets, or long-lived connections. Wait for the data selector or response instead.
Cloud Run work pauses after returning a response CPU is disabled after the response. Configure CPU to remain allocated while Puppeteer work runs.
Browser crashes under load Memory pressure, leaked pages, or excessive parallelism. Close targets in finally, cap workers, monitor memory, and recycle workers when appropriate.

12. A practical tuning checklist

  • Define the exact record-ready condition for each site.
  • Replace fixed sleeps with selector, request, response, or navigation waits.
  • Use domcontentloaded when initial HTML is sufficient.
  • Block only verified-unneeded resource types.
  • Reuse one browser process and isolate jobs with pages or contexts.
  • Start with a small worker count and increase gradually.
  • Measure p50, p95, timeout rate, bytes, CPU, memory, and records per minute.
  • Keep CPU allocated for background browser work in serverless deployments.
  • Retry transient failures with limits and backoff.
  • Check permission, privacy, robots, terms, and rate-limit requirements before scaling.

13. Or skip the browser setup

If your goal is reliable page images rather than DOM extraction, ScreenshotNeo provides a single screenshot API request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

14. Cost and performance notes

Browser speed is a capacity-planning problem. More pages increase throughput only while CPU, memory, bandwidth, and the target’s allowed rate have headroom. Blocking assets can reduce work but can also change page behavior. Network-idle waits can improve correctness for some applications while making others much slower. Measure the complete pipeline and the records returned, not just navigation time.

Puppeteer’s installation documentation lists approximate Chrome for Testing download sizes of 170 MB on macOS, 282 MB on Linux, and 280 MB on Windows. Those are download sizes, not scraping-speed benchmarks. The official sources do not establish a universal percentage or “10x” speedup.

FAQ

Should I always block images?

No. Block them when your extracted data does not depend on image-triggered code, layout, or lazy-loading behavior. Verify each site.

Is one browser per URL faster?

Usually no. Browser startup and memory overhead make a reused browser with bounded pages or contexts more efficient.

How many Puppeteer pages can run at once?

There is no fixed number. Start small, measure, and increase until a resource limit, error rate, or permitted target rate becomes the bottleneck.

When is networkidle appropriate?

Use it when network quiescence itself proves readiness. Prefer a selector or response when a specific data condition is available.

Can I scrape faster by bypassing a bot check?

Do not treat bypassing access controls as a performance technique. Follow the site’s permission, terms, and rate-limit requirements.