ScreenshotNeo

BlogHow-to

How to Automate a Browser with Puppeteer

Learn how to install Puppeteer, navigate pages, click buttons, wait for app state, extract data, and save screenshots or PDFs with reliable cleanup.

By the ScreenshotNeo team4 October 20269 min read

Puppeteer automates Chrome or Firefox from JavaScript. The basic workflow is to launch a browser, open a page, navigate to a URL, interact with elements using locators, wait for the page state your task needs, collect or save the result, and close the browser in a finally block.

This guide shows how to install Puppeteer, click a button, take a screenshot, handle single-page apps, choose a browser, and diagnose common failures. Puppeteer runs headless by default; you can configure a visible browser when you need to watch the interaction. Puppeteer: What is Puppeteer?

1. Install Puppeteer and prepare a project

Use a supported Node.js version for the Puppeteer release you install. The current documentation snapshot specifies Node 22.12 or newer; check the system requirements for the release you choose. puppeteer downloads a compatible browser during installation. Use puppeteer-core when you manage or connect to a browser separately.

mkdir puppeteer-task
cd puppeteer-task
npm init -y
npm install puppeteer

Set the package to use ECMAScript modules by adding "type": "module" to package.json, or save the example as .mjs. Create capture.js:

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1280, height: 800 });
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  console.log(await page.title());
} finally {
  await browser.close();
}

Run it with node capture.js. Puppeteer pairs releases with browser versions, so consult the supported browsers table rather than assuming an arbitrary installed browser will match. The version numbers in that table change over time.

2. How do I automate a browser with Puppeteer?

The essential sequence is launch → page → navigate → interact → capture or extract → close. This complete example searches a page, reads a result, saves a screenshot, and guarantees browser cleanup even if navigation or an action fails. Change the URL and selectors to match the site you control or are authorized to automate.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1365, height: 900 });

  const response = await page.goto('https://example.com', {
    waitUntil: 'domcontentloaded',
    timeout: 30_000,
  });
  if (!response) {
    throw new Error('Navigation did not produce a main-resource response');
  }
  if (!response.ok()) {
    throw new Error(`Page returned HTTP ${response.status()}`);
  }

  // Locators wait for the target to be ready before acting.
  await page.locator('input[name="q"]').fill('Puppeteer');
  await page.locator('button[type="submit"]').click();

  // Wait for the state your task actually needs.
  await page.locator('[data-test="result-title"]').wait();
  const title = await page.locator('[data-test="result-title"]').evaluate(el => el.textContent?.trim());
  console.log('Result:', title);

  await page.screenshot({ path: 'result.png', fullPage: true });
} finally {
  await browser.close();
}

The sample selectors are illustrative: a real site may use different fields and result markup. Prefer a stable test attribute, accessible name, or text target where available instead of relying on a long chain of layout-dependent CSS selectors. Puppeteer recommends locators for ordinary interactions because they wait for the target and check action readiness. Page interactions guide

3. How do I click a button with Puppeteer?

Use a locator and call click(). Locators wait for an element to exist and be suitable for the action, including checks related to visibility, enabled state, viewport presence, and a stable bounding box. They retry when the target is not ready. For example:

await page.locator('button[type="submit"]').click();
await page.locator('input[name="email"]').fill('dev@example.com');

CSS selectors work by default. Puppeteer also supports selector syntax for text, accessibility attributes, XPath, and Shadow DOM. For a label or accessible name, use a selector such as ::-p-aria(Submit); for visible text, use ::-p-text(Continue). Confirm the target uniquely identifies the intended control.

Use page.keyboard, page.mouse, or page.touchscreen when the task requires direct input events rather than a locator action. Lower-level waitForSelector() and element handles are available, but a selector wait does not automatically retry a later action. Dispose handles you keep around when finished to avoid retaining page objects unnecessarily.

4. How do I take a screenshot with Puppeteer?

Navigate to the page, wait for the content that should appear in the image, then call page.screenshot(). This saves a full-page PNG:

await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.locator('main').wait();
await page.screenshot({ path: 'page.png', fullPage: true });

For a viewport-only image, omit fullPage. To capture one element, use an element handle’s screenshot() method; Puppeteer can scroll it into view first:

await page.locator('article').screenshot({ path: 'article.png' });

Screenshot options include a path, image type (PNG, JPEG, or WebP where supported by the installed release), quality for lossy formats, full-page capture, clipping, and transparent background with omitBackground: true. The image type can be inferred from the file extension. Check the installed release’s ScreenshotOptions API for the exact supported options. Full-page images can be very tall, and lazy-loaded content may not appear until it has been brought into view or otherwise loaded; wait for the page’s own content state before capture.

5. Wait for the right state, especially in single-page apps

Navigation waiting and application readiness are different. Puppeteer treats URL changes as navigation, including anchor changes and History API changes, which is useful for single-page apps. A changed URL does not prove that the data or component your task needs has rendered. After navigation or a click, wait for the specific result element, text, or state before reading or capturing it. Puppeteer FAQ

page.goto() accepts navigation wait conditions such as load, domcontentloaded, and networkidle0/networkidle2. Choose according to the task: domcontentloaded is often a useful initial milestone, while a selector wait expresses the actual application condition. Network-idle conditions can be a poor fit for pages with long polling, analytics, or persistent connections. Avoid arbitrary fixed sleeps when a concrete condition is available.

await page.goto('https://example.com/app', { waitUntil: 'domcontentloaded' });
await page.locator('[data-test="account-summary"]').wait();
const summary = await page.locator('[data-test="account-summary"]').evaluate(el => el.textContent?.trim());

6. Choose Chrome, Firefox, headless, or visible mode

Puppeteer supports Chrome and Firefox from v23.0.0. It uses CDP by default for Chrome and WebDriver BiDi by default for Firefox. Both protocols have production-ready support, but their feature coverage differs. Keep the browser and Puppeteer versions aligned and consult the protocol documentation if your script depends on browser-specific capabilities. Cross-browser support FAQ

Headless is the default. To inspect a run in a visible browser, launch with headless: false:

const browser = await puppeteer.launch({ headless: false });

For a separately installed browser, puppeteer-core is the smaller control package; provide the executable or connect to a browser endpoint as appropriate. Puppeteer’s browser management tooling can install stable or pinned Chrome for Testing versions. Its installation requirements differ by platform; on Linux and macOS, Chrome extraction may require unzip, while Windows uses tar.exe or PowerShell. See browser management documentation and system requirements.

7. Generate a PDF

Use page.pdf() after navigation and any required readiness wait. PDF generation uses print CSS media by default. If the document should use screen styles, emulate screen media first:

await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.locator('main').wait();
await page.emulateMediaType('screen');
await page.pdf({ path: 'report.pdf', format: 'A4', printBackground: true });

PDF options include format, dimensions, margins, landscape layout, background printing, page ranges, and headers/footers. Consult the Page.pdf API for the installed version’s complete option list.

8. Reliability, performance, and operating cost

  • Always close resources. Put browser.close() in finally. If reusing a browser for multiple tasks, close each page or context when done and still close the browser at process shutdown.
  • Use explicit timeouts and meaningful waits. A navigation timeout can bound a stalled page; locator waits should target the element or state required. Handle expected missing elements explicitly rather than silently scraping stale or empty content.
  • Limit concurrency. Each page consumes browser and system resources. Start with a small number of concurrent pages, measure memory and duration in your environment, and increase gradually. No universal throughput figure applies across sites and machines.
  • Keep captures bounded. Full-page screenshots, high device scale factors, many open pages, and large PDFs use more memory and time than a viewport capture. Capture only the area and output needed.
  • Separate browser startup from task time. For repeated jobs, a managed long-lived browser can avoid launching for every URL, but isolate sessions with separate browser contexts when cookies and storage must not carry over.
  • Budget for infrastructure. Puppeteer itself is an open-source library; operating cost comes from the machine, browser runtime, storage, and any network or proxy services your workflow requires. Measure actual workloads rather than relying on undocumented benchmarks.

9. Troubleshooting common Puppeteer errors

Symptom Likely cause Fix
Browser executable not found Using puppeteer-core without configuring a browser, or the downloaded browser is missing. Install puppeteer for its managed browser, or install a compatible browser and configure the executable/connection. Check the version support table.
Browser fails to launch in Linux or a container Missing system libraries, sandbox constraints, or browser extraction utilities. Install the documented platform dependencies and verify the container supports the browser. Avoid disabling the sandbox as a default workaround; use a secure runtime configuration appropriate to your deployment.
Timeout waiting for navigation The page is slow, never reaches the selected lifecycle event, or keeps network activity open. Choose a more appropriate waitUntil, set a justified timeout, and separately wait for the task’s target element.
Click times out or hits the wrong control The selector is ambiguous, the control is hidden/disabled, an overlay blocks it, or the layout changed. Use a unique accessible, text, or stable test selector. Wait for the expected state and inspect the page in visible mode when debugging.
Script reads empty or old SPA content The URL changed before the application finished rendering its data. Wait for the result element or a specific text/state after the URL transition.
Screenshot misses images or lower-page content Lazy content has not loaded, the screenshot is viewport-only, or the page needs scrolling. Use fullPage: true when appropriate and wait for or trigger lazy content loading before capture.
PDF looks different from the browser PDF uses print media CSS by default. Call page.emulateMediaType('screen') before page.pdf() if screen styling is required.
Process memory grows over time Pages, browser instances, or element handles remain open or referenced. Close pages/contexts, dispose retained handles, and close the browser in cleanup code.

10. Or skip the browser setup

If the job is simply to capture a URL, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. One GET request returns an image or PDF; see the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month, no card required.

11. Frequently asked questions

Does Puppeteer work with Firefox?

Yes. Puppeteer supports Firefox from v23.0.0 and uses WebDriver BiDi by default for Firefox. Check feature support for the protocol and release you plan to use.

Can Puppeteer automate a site that requires login?

It can interact with login forms and browser storage, subject to the site’s rules and your authorization. Keep credentials out of source code and logs, and use isolated browser contexts when separate sessions are required.

Can Puppeteer run without a visible desktop?

Yes. Headless mode is the default, so it can run in a server or CI environment that has the required browser dependencies.

Should I use Puppeteer or an API for one screenshot?

Use Puppeteer when the task needs custom browser logic, test assertions, or a sequence of interactions. A screenshot API is simpler when the input is a URL and the desired output is an image or PDF.