ScreenshotNeo

BlogEngineering

Puppeteer Benchmarking: How to Measure Browser Automation Performance

Build repeatable Puppeteer benchmarks that separate automation latency, browser runtime metrics, and page performance, with runnable code and reporting guidance.

By the ScreenshotNeo team4 October 20269 min read

To benchmark Puppeteer, define a repeatable browser task, pin the browser and runtime conditions, and measure the task end to end across repeated runs. Use Puppeteer metrics and Chrome traces to diagnose browser work, and use Lighthouse separately when the question is page performance. These are different measurements; a Lighthouse score is not Puppeteer execution speed.

Puppeteer is a JavaScript library for controlling Chrome or Firefox through the DevTools Protocol or WebDriver BiDi. Its APIs include page metrics and tracing. The key to a useful benchmark is to say exactly what was timed, under which conditions, and how failures and run-to-run variation were handled. See the Puppeteer overview and Page API documentation.

1. Choose the question and timing boundary

Start by deciding what decision the benchmark should inform. “How fast is Puppeteer?” is too broad: launching a browser, loading a page, waiting for an element, clicking through a flow, and extracting data all contribute differently.

Question Measure Include in the result
How long does this automated task take? End-to-end task duration State whether launch, navigation, waits, actions, extraction, and teardown are inside the timed boundary.
What browser work explains a slow run? Runtime metrics and a trace Use metrics and trace events for diagnosis; collect traces on representative runs.
How does the page itself perform? Lighthouse audit Report page audit results separately from automation duration.

For production automation, define a realistic task such as navigating to a stable page, waiting for a known element, performing an interaction, and extracting a result. Decide whether the browser is cold-started for every run or reused. If both situations matter, benchmark them as separate scenarios.

2. Make the workload reproducible

  1. Choose a stable target page or fixture, test data, and a specific interaction path. Control network conditions where possible, or record their variation.
  2. Pin and record Node.js, Puppeteer, browser build, operating system, machine resources, browser mode, and protocol.
  3. Keep compared configurations equivalent. If comparing CDP and BiDi, use the same task and stratify by platform and browser mode.
  4. Choose and document a warm-up policy, then perform repeated measured runs. Keep each run’s duration, outcome, timeout status, and relevant metrics.
  5. Summarize the distribution, such as median and percentiles or interquartile range. Do not report only the fastest run. Retain failures and timeouts as outcomes.

These reporting practices help readers interpret the result; Puppeteer does not prescribe a universal run count, pass threshold, or performance target. Set a workload-specific baseline and target. Chromium’s BiDi benchmark page illustrates why protocol, operating system, and browser mode should be kept distinct, and flags known flakiness in some Mac comparisons.

3. Runnable Puppeteer benchmark with metrics and tracing

The following Node.js script times a complete navigation-and-selector task, records Puppeteer page metrics, and saves a trace for inspection. It runs repeated trials in one browser process, opening a fresh page per trial. Set TARGET_URL to a stable page you are authorized to access. The trace is diagnostic instrumentation; because tracing adds overhead, this example does not use its traced duration as the uninstrumented latency result.

const puppeteer = require('puppeteer');

const targetUrl = process.env.TARGET_URL || 'https://example.com';
const runs = Number(process.env.RUNS || 10);
const warmups = Number(process.env.WARMUPS || 2);
const selector = process.env.SELECTOR || 'h1';

function percentile(values, p) {
  const sorted = [...values].sort((a, b) => a - b);
  const index = Math.min(sorted.length - 1, Math.ceil(p * sorted.length) - 1);
  return sorted[index];
}

async function oneRun(browser, index, saveTrace = false) {
  const page = await browser.newPage();
  let tracePath;
  try {
    if (saveTrace) {
      tracePath = `trace-${index}.json`;
      await page.tracing.start({ path: tracePath, screenshots: false });
    }
    const start = process.hrtime.bigint();
    await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 30000 });
    await page.waitForSelector(selector, { timeout: 10000 });
    const result = await page.$eval(selector, el => el.textContent.trim());
    const durationMs = Number(process.hrtime.bigint() - start) / 1e6;
    const metrics = await page.metrics();
    return { ok: true, durationMs, result, metrics, tracePath };
  } catch (error) {
    return { ok: false, error: error.message, tracePath };
  } finally {
    if (saveTrace) {
      try { await page.tracing.stop(); } catch {}
    }
    await page.close();
  }
}

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    for (let i = 0; i < warmups; i++) await oneRun(browser, `warmup-${i}`);
    const results = [];
    for (let i = 0; i < runs; i++) {
      // Collect a diagnostic trace on the first measured run only.
      results.push(await oneRun(browser, i, i === 0));
    }
    const successful = results.filter(r => r.ok);
    const durations = successful.map(r => r.durationMs);
    console.log(JSON.stringify({
      targetUrl, runs, warmups,
      successes: successful.length,
      failures: results.filter(r => !r.ok),
      latencyMs: durations.length ? {
        median: percentile(durations, 0.5),
        p90: percentile(durations, 0.9),
        min: Math.min(...durations),
        max: Math.max(...durations)
      } : null,
      perRun: results
    }, null, 2));
  } finally {
    await browser.close();
  }
})().catch(error => { console.error(error); process.exitCode = 1; });

Install Puppeteer in a project with npm install puppeteer, save the script as benchmark.js, then run TARGET_URL=https://example.com RUNS=20 WARMUPS=3 node benchmark.js. The default target is only a runnable example. For meaningful comparisons, use your actual workflow or a controlled fixture.

The timer starts immediately before navigation and ends after the selector is found and text is extracted. It excludes browser launch, page creation, metric retrieval, trace stopping, and teardown. If cold-start time matters, add a separate timer around puppeteer.launch(); do not silently fold launch time into only some configurations.

4. Interpret metrics and inspect traces

page.metrics() returns browser runtime counters exposed by Puppeteer. They can help explain a run, but they are not a standalone score and should be interpreted alongside the task and environment. Puppeteer tracing records a timeline that can be opened in Chrome DevTools or a compatible timeline viewer. Capture traces for typical or slow runs, and keep trace collection out of the primary latency measurement unless trace overhead is intentionally part of the workload.

For lower-level Chrome runtime metrics, the Chrome DevTools Protocol Performance domain supports enabling collection and retrieving current runtime metrics. Consult the Performance domain reference for the protocol methods and returned data. Protocol-level collection is useful when you need a more explicit CDP instrumentation path; it does not replace a clearly defined end-to-end task timer.

5. Use Lighthouse for page audits, not automation timing

When the goal is to assess page loading or runtime performance, run Lighthouse under a declared configuration. Lighthouse can be used programmatically as a Node module and can audit public or authenticated pages. Keep its results in a separate section or dataset from Puppeteer task timings: Lighthouse audits page characteristics, while your Puppeteer timer measures the chosen automation workflow. See the Lighthouse documentation.

6. Compare protocols, platforms, and browser modes fairly

  • Protocol: Compare CDP and WebDriver BiDi only when both setups support the same scenario and are configured equivalently.
  • Operating system: Report platforms separately or show each stratum; do not pool different machines into one unexplained number.
  • Browser mode: Keep headless shell, newer headless, and headful runs distinct when those modes are relevant.
  • Reliability: Publish successes, failures, timeouts, and known flaky conditions next to latency.
  • Measurement layer: Do not combine task duration, runtime counters, trace observations, and Lighthouse results as if they measured the same thing.

The Chromium BiDi benchmark results are comparative evidence for specific configurations, not a universal speed ranking for every Puppeteer workload. There is no general pass/fail number in the cited documentation that applies to all tasks.

7. Publish a benchmark others can reproduce

Include these details with the result:

  • Scenario steps, target or fixture description, and timing boundaries.
  • Exact Node.js, Puppeteer, and browser versions, plus OS, machine resources, browser mode, and protocol.
  • Network and test-data conditions, warm-up policy, and number of measured runs.
  • Per-run data or a usable summary with median and spread, alongside failures and timeouts.
  • Whether browser startup and teardown were measured, and whether tracing was enabled.
  • Trace artifacts or raw data for relevant representative runs, where sharing is appropriate.

8. Common benchmark problems

Symptom Likely cause Fix
One run looks dramatically faster Warm caches, startup variance, background machine load, or network variation. Repeat runs, declare warm-up behavior, preserve per-run values, and report distribution and environment.
Comparisons change when traces are enabled Tracing adds instrumentation overhead. Keep primary latency runs untraced; use separate representative traced runs for diagnosis.
Navigation times out The page did not reach the chosen navigation condition before the timeout, or the target is unstable. Choose a condition that matches the task, use a controlled fixture where possible, and record timeouts instead of silently dropping them.
Selector wait times out The selector is absent, appears conditionally, or the page state differs across runs. Verify the selector and scenario preconditions; retain the failed run and error in the report.
Different machines produce conflicting results OS, CPU, memory, browser build, mode, or protocol differs. Pin the environment or present separate results by configuration.
Lighthouse result is mistaken for Puppeteer speed Page audit output and automation task timing are being conflated. Report them separately and state the question each measurement answers.
Metric values look unrelated to task latency Runtime counters describe browser activity, not elapsed workflow time. Use them as diagnostic context and retain the end-to-end timer as its own measure.

9. Performance, reliability, and cost considerations

Browser startup can dominate short tasks, so measure both cold and reused-browser cases if both occur in production. Network-dependent pages add variance; a controlled fixture improves repeatability, while a live-site scenario may better represent deployed conditions. These should be separate test cases where each answers a different operational question.

Reliability belongs in the benchmark result: a configuration that is fast only on successful runs can mislead. Report failure and timeout counts with latency, and investigate traces for representative slow runs. Avoid treating a single best time as a guarantee or extrapolating one page’s result to other workloads.

Benchmarking has resource costs: launching browsers, repeating navigation, retaining traces, and running Lighthouse all consume machine time and may generate network traffic. Bound the run count to the decision being made, avoid unnecessary repeated tests against third-party production sites, and store only the trace and raw data needed to support the conclusion. The research sources do not establish a universal monetary cost or runtime for a Puppeteer benchmark; these depend on the target, environment, and workload.

Or skip the browser setup

If your benchmark task is capturing website screenshots, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. It returns PNG, JPEG, WebP, or PDF. Its response headers identify the page verdict and billing status; bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits cost nothing. Cookie consent banners are accepted and removed, and known consent platforms, newsletter popups, and chat widgets can be removed before capture; each cleanup step can be turned off.

Use the DIY Puppeteer workflow above when you need to benchmark browser automation itself. For a screenshot workload, this one call avoids managing browser launch and capture code in your own script. It is still a different measurement: benchmark the API request separately if you are comparing screenshot services.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. The MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card.

FAQ

How many benchmark runs should I do?

There is no universal count for every workload. Use enough repeated runs to see the variation relevant to your decision, and publish the count and per-run outcomes.

Should browser launch be included?

Include it when cold-start behavior is the question. Otherwise measure launch separately and state that the task timer begins with an already running browser.

Can Puppeteer traces tell me why a task is slow?

They can expose browser timeline activity for diagnosis. Pair trace inspection with the task’s defined timing and avoid using instrumented runs as unqualified latency figures.

Does a faster benchmark mean a faster user experience?

Not by itself. A benchmark describes its defined task and environment; use page audits and user-facing measurements when those are the question.