ScreenshotNeo

BlogEngineering

Puppeteer Benchmark Tools: How to Measure Browser Automation Performance

Measure Puppeteer performance with repeatable workloads, clear timing boundaries, and separate reports for speed, reliability, and resource use.

By the ScreenshotNeo team4 October 202610 min read

To measure Puppeteer performance, run the same browser automation workload repeatedly under controlled conditions. Record elapsed time and throughput separately from success rate and resource use. Report the Puppeteer and browser versions, protocol, operating system, browser mode, machine, and network or CPU conditions so another developer can interpret or reproduce the result.

A trace helps explain where time went; it is not a benchmark result on its own. Lighthouse audits page performance, while a Puppeteer benchmark measures the execution of your automation workflow. Keep those results separate.

1. Decide what you are measuring

Choose one question before writing the benchmark. Different questions need different timing boundaries and should not be combined into one result.

Question Example timing boundary Useful companion data
Browser launch cost Immediately before launch to browser ready Cold or warm start, browser process count
Navigation cost Before navigation to a defined load condition Navigation errors, selected wait condition
User-like workflow Before first action to a meaningful completion state Action-level timings and failure reason
Extraction Before extraction to validated data Rows or records returned, validation result
Screenshot generation Before capture to image file available Viewport, full-page setting, image format
Sustained concurrency Fixed load interval or number of completed jobs Concurrency, queue time, throughput, failures

Define a success condition, such as a selector appearing and extracted data passing validation. A run that ends quickly because navigation failed is not a performance improvement. Fix the URL, data, action sequence, timeout policy, browser settings, and completion condition across the candidates being compared.

2. Fix the environment and record it

Browser automation timings can change with the machine, browser build, protocol, and mode. Record enough context to make the comparison meaningful:

  • Puppeteer version and exact browser version.
  • Protocol in use: Chrome DevTools Protocol (CDP) or WebDriver BiDi.
  • Operating system and version, machine CPU and memory class, and whether other load was present.
  • Headless, headful, or headless-shell mode, plus viewport and device scale factor.
  • Network and CPU throttling settings, if any.
  • Workload URL or fixture, test data, timeout policy, and completion condition.
  • Whether each run is a cold start or reuses a browser/page.
  • Concurrency and the method used to collect CPU and memory.

Do not compare results collected on different browser builds or machine loads as if the only variable were Puppeteer. If protocol or browser mode is the variable of interest, keep the other conditions fixed and state the exact configuration.

3. Build a repeatable benchmark

The following runnable Node.js example measures a fixed workflow: launch a browser, navigate, wait for a known selector, read a title, and close. It records every successful duration and failure, reports median and p95 for successful runs, and keeps reliability visible. Set BENCH_URL and BENCH_SELECTOR for a stable page you are authorized to automate.

const puppeteer = require('puppeteer');

const url = process.env.BENCH_URL || 'https://example.com';
const selector = process.env.BENCH_SELECTOR || 'h1';
const runs = Number(process.env.BENCH_RUNS || 20);

function percentile(sorted, p) {
  if (sorted.length === 0) return null;
  return sorted[Math.min(sorted.length - 1, Math.ceil(p * sorted.length) - 1)];
}

async function once() {
  const started = performance.now();
  let browser;
  try {
    browser = await puppeteer.launch({ headless: true });
    const page = await browser.newPage();
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
    await page.waitForSelector(selector, { timeout: 15000 });
    const title = await page.title();
    if (!title) throw new Error('Validation failed: page title is empty');
    return { ok: true, ms: performance.now() - started };
  } catch (error) {
    return { ok: false, ms: performance.now() - started, error: error.message };
  } finally {
    if (browser) await browser.close();
  }
}

(async () => {
  const samples = [];
  for (let i = 0; i < runs; i++) {
    samples.push(await once());
  }
  const good = samples.filter(x => x.ok).map(x => x.ms).sort((a, b) => a - b);
  const failures = samples.filter(x => !x.ok);
  console.log(JSON.stringify({
    url, selector, runs, successes: good.length, failures: failures.length,
    successRate: good.length / runs,
    medianMs: percentile(good, 0.50), p95Ms: percentile(good, 0.95),
    errors: failures.map(x => x.error), samples
  }, null, 2));
})().catch(error => { console.error(error); process.exitCode = 1; });

Save it as bench.js, install Puppeteer in the project with npm install puppeteer, then run BENCH_RUNS=30 BENCH_URL=https://your-test-site.example BENCH_SELECTOR='main h1' node bench.js. The first run may include browser download or setup work; install dependencies before collecting measurements. For production-like tests, prefer a stable staging page or fixture over a public page that can change, throttle, or block automation.

This example intentionally includes launch time in each sample. To measure a reused browser instead, launch once outside the loop and create a fresh page for each run; state that choice in the report. To measure only a workflow inside an already-open page, start the timer immediately before the first workflow action and stop it at the completion condition. Keep launch, navigation, action, extraction, and capture timings separate when they answer different operational questions.

4. Measure latency, throughput, reliability, and resources

Latency

Latency is elapsed time to a named milestone or for the complete task. Use a monotonic timer such as performance.now(). Report the number of repetitions, median, and a spread such as p95 or interquartile range. Include raw samples when practical. A mean alone can hide a few slow runs; a median alone can hide tail delays.

Throughput

Throughput is completed tasks per unit time at a stated concurrency. Report the number of workers or simultaneous pages, test duration or task count, and whether queueing time is included. A single-run latency does not establish concurrent throughput. Increase concurrency in controlled steps and record failures and resource use at each step.

Reliability

Report successes, failures, timeouts, and error categories alongside timing. Define whether retries count as part of the user-visible task and whether a retried task is considered successful. Do not silently discard failed or unusually slow runs; explain exclusions and preserve the raw samples.

Resource use

Measure CPU and memory with the same operating-system or container monitoring method for every candidate. State the sampling interval and whether measurements include the browser process tree, the Node.js process, or both. For sustained runs, also note browser process count and whether resource use grows over time. These figures need the same environmental controls as timing to support a fair comparison.

5. Repeat runs and summarize uncertainty

  1. Run a small warm-up if startup, JIT, or cache state could affect the measurement. Label warm-up runs and do not mix them into the reported sample unless warm-up is part of the real workload.
  2. Collect enough repetitions to show variation. Choose a run count before comparing candidates, and use the same count and stopping rules for each.
  3. Keep cold-start and warm runs in separate groups. Do not combine a first launch with reused-browser timings.
  4. Report sample count, median, spread or percentiles, success rate, and failure categories. Preserve the raw timing samples and any trace files.
  5. If reporting confidence intervals, describe the calculation and the measured unit. Do not infer a universal winner from a small or configuration-specific run.

For comparisons, change one factor at a time where possible: protocol, browser version, operating system, headless/headful mode, workload, or concurrency. The Chromium BiDi benchmark project illustrates why these dimensions matter: its index compares configurations and reports relative overhead with 95% confidence intervals, while flagging known flakiness in some Mac comparisons. Treat those results as specific to their configurations rather than a general ranking.

6. Use profiling and audit tools for the right question

Puppeteer tracing

Puppeteer can capture a timeline trace to help diagnose performance issues. A trace is diagnostic evidence: inspect it to see activity and timing, but do not substitute its contents for repeated end-to-end measurements. The official [Puppeteer guide](https://pptr.dev/guides/diagnostics) describes tracing, and the [tracing API](https://pptr.dev/api/puppeteer.tracing) can start and stop tracing to create a file that opens in Chrome DevTools or another timeline viewer.

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.tracing.start({ path: 'trace.json' });
    await page.goto('https://example.com', { waitUntil: 'networkidle2' });
    await page.waitForSelector('h1');
    await page.tracing.stop();
  } finally {
    await browser.close();
  }
})().catch(error => { console.error(error); process.exitCode = 1; });

Use the same workload and conditions when capturing comparison traces. Tracing adds work and can affect timings, so collect baseline timing runs separately unless trace overhead itself is what you are measuring.

Chrome DevTools Performance panel

Use the Performance panel to inspect CPU profiles and runtime behavior after or during a targeted diagnostic session. It can show local Core Web Vitals such as LCP, CLS, and INP, and offers CPU and network throttling controls. Advanced instrumentation can significantly hinder performance; leave it off during baseline timing unless instrumentation overhead is the subject of the experiment. See the [Chrome DevTools Performance documentation](https://developer.chrome.com/docs/devtools/performance/).

Lighthouse

Lighthouse audits page performance and related quality areas. Puppeteer can hand a page or browser off to Lighthouse, or connect to a browser instance Lighthouse launched, as described in the [Lighthouse Puppeteer integration guidance](https://github.com/GoogleChrome/lighthouse/blob/main/docs/puppeteer.md). Keep Lighthouse audit scores separate from Puppeteer script latency and throughput: they answer different questions. See the [Lighthouse documentation](https://developer.chrome.com/docs/lighthouse/).

7. Compare protocols and execution modes fairly

If comparing CDP with WebDriver BiDi, use the same task, browser build where supported, machine, page state, wait conditions, and completion checks. Record the protocol and any differences in capabilities that force a workflow change. Also compare cold versus warm execution, headless versus headful mode, operating system, workload size, and concurrency as separate axes. The [Chromium BiDi benchmark index](https://chromium.googlesource.com/chromium/src/+/main/chrome/test/chromedriver/bidi/benchmarks/README.md) is useful as an example of configuration-aware reporting; its indexed results do not establish a universal protocol winner.

8. Troubleshooting common benchmark problems

Symptom Likely cause Fix
Large variation between identical runs Machine load, network changes, cache state, or variable page content Use a controlled fixture, isolate cold and warm runs, record machine load, and repeat with a fixed workload.
Very fast runs mixed with failures Navigation or selector waits are failing early Require a meaningful completion condition, log errors, and report failures with timing.
Timeouts on pages with ongoing requests A network-idle wait condition may never be satisfied by analytics, streaming, or polling Use a task-specific selector or completion signal, and use the same wait condition for each candidate.
First sample is much slower Cold browser launch, cache population, or setup work is included Report cold start separately from warm runs, and ensure dependencies are installed before measurement.
Trace run is slower than baseline Tracing or profiling adds overhead Collect clean timing runs without diagnostic instrumentation; use traces to explain behavior in a separate pass.
Different results between CDP and BiDi Protocol, browser support, or configuration differs Record exact versions and modes; compare the same workflow and report any capability-specific changes.
Throughput stops rising or failures increase under load CPU, memory, target-site limits, or browser process capacity is saturated Record resource use and failures at each concurrency level; lower concurrency or allocate more capacity before drawing conclusions.
Benchmark results disagree with Lighthouse Automation speed and page audit metrics are different measurements Keep task latency/throughput and Lighthouse audit results in separate reports.

9. Performance, reliability, and cost considerations

For an apples-to-apples performance result, include all work required by the actual task: launch if the deployed system launches per job, navigation, waits, actions, extraction, and output writing if those are user-visible. Exclude setup that happens only once in production only when the reported metric also excludes it. Reusing a browser can reduce per-task overhead, but it changes the operational model and may introduce state leakage; use a fresh page or context when the real system needs isolation.

Reliability is part of capacity planning. A faster configuration that times out more often may deliver fewer successful tasks. For a sustained run, track successful completions per time interval along with errors and resource growth. Avoid running a high-concurrency benchmark against a third-party site without authorization; use a controlled target to prevent rate limits and variable external behavior from dominating the result.

Cost depends on the runtime environment and workload, not on a timing number alone. For self-hosted automation, account for browser CPU and memory, execution duration, concurrency, retries, and infrastructure idle time. For a screenshot API, compare the billable event definition, failure handling, cache behavior, and required features as well as the per-request price. Do not extrapolate cost from a throughput result without stating those assumptions.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF screenshot. Its capture flow accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. The same parameters used by other screenshot APIs also work, which can make switching straightforward. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Is a trace enough to benchmark Puppeteer?

No. A trace is for diagnosing timeline behavior. Benchmark results come from repeated runs with defined timing boundaries and reported outcomes.

Should I use Lighthouse to measure script speed?

No. Lighthouse audits page performance. Use a controlled Puppeteer workload for automation latency and throughput, and report audit results separately.

How many repetitions are enough?

There is no fixed count that fits every workload. Choose a count that exposes variability, apply it consistently, and report the sample count, spread, and failures.

Can I compare headless and headful results?

Yes, if you hold the other conditions constant and label each mode. Treat them as separate configurations because they can behave differently.