ScreenshotNeo

BlogGuides

Puppeteer Web Scraping: The Complete Guide

Learn how to scrape JavaScript-rendered pages with Puppeteer, wait reliably, extract and validate data, and troubleshoot common browser automation failures.

By the ScreenshotNeo team4 October 20269 min read

Use Puppeteer when the content you need depends on browser-side JavaScript or an interaction such as clicking a button. It controls Chrome or Firefox so your code can navigate, wait for a page state, interact with elements, and read the resulting DOM. It is not necessary for every website, and using it does not grant permission to collect a site’s data.

This guide builds a small JavaScript scraper, explains browser setup and reliable waits, and covers extraction, validation, screenshots, PDFs, failures, and responsible operation.

1. What Puppeteer does—and when to use it

Puppeteer is a JavaScript library for controlling Chrome or Firefox through the DevTools Protocol or WebDriver BiDi. It runs headless by default. A scraper built with Puppeteer can access content after the browser executes scripts, and can perform actions that make more content appear.

Before choosing a browser, check whether the required data is already available in the initial HTML or through a documented data endpoint you are permitted to use. A browser adds startup, memory, and page-rendering work, so it is useful when rendering or interaction is part of the task—not as a requirement for all web pages.

2. Install Puppeteer and choose a browser setup

The two common package choices are puppeteer and puppeteer-core:

Package Browser setup Choose it when
puppeteer Downloads a compatible Chrome during installation. You want the package to manage the browser download.
puppeteer-core Installs the library without downloading Chrome. Your environment provides and configures the browser separately.

For a typical project using npm and the package-managed browser:

npm install puppeteer

If your deployment environment manages its own browser, install puppeteer-core instead and configure the executable path or connection as appropriate for that environment:

npm install puppeteer-core

Some package managers or deployment setups block install scripts. If that prevents the browser download, follow the current official installation instructions; Puppeteer documents npx puppeteer browsers install as a manual installation route. Browser versions change, so use the documentation for the version you install rather than pinning an assumed current version.

3. A complete scraper in JavaScript

This runnable example launches a browser, opens a page, waits for a specific result, extracts and validates its text, and closes the browser even if navigation or extraction fails. Replace the example URL and selector with a page and element you are allowed to access.

const puppeteer = require('puppeteer');

async function main() {
  const browser = await puppeteer.launch({ headless: true });

  try {
    const page = await browser.newPage();
    page.setDefaultTimeout(15_000);
    page.setDefaultNavigationTimeout(30_000);

    const response = await page.goto('https://example.com/', {
      waitUntil: 'domcontentloaded',
      timeout: 30_000,
    });

    if (!response) {
      throw new Error('Navigation did not return a main-resource response');
    }
    if (!response.ok()) {
      throw new Error(`Page returned HTTP ${response.status()}`);
    }

    const result = page.locator('h1');
    await result.wait();
    const title = (await result.map(el => el.textContent).wait()).trim();

    if (!title) {
      throw new Error('The matching heading was empty');
    }

    console.log({ url: page.url(), title });
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The example uses CommonJS syntax. In a project configured for ECMAScript modules, use import puppeteer from 'puppeteer'; and keep the same asynchronous flow. The target page’s structure determines the selector and the right readiness condition.

4. Navigate and wait for the page state you need

A completed navigation call does not prove that the data you want is present. Choose a wait that matches the page and task, then confirm the expected element or value.

Need Approach Tradeoff
Initial document parsed waitUntil: 'domcontentloaded' Does not mean all resources or app-rendered data are ready.
Load event fired waitUntil: 'load' Some pages keep loading resources after useful content appears.
Network mostly idle waitUntil: 'networkidle0' or 'networkidle2' Persistent connections or background requests can make network-idle waits unsuitable.
A particular result exists page.locator(selector).wait() or page.waitForSelector(selector) Usually the clearest choice when the desired content has a stable selector.
A particular response arrives page.waitForResponse(predicate) Useful when the page fetches the needed data after navigation; inspect and handle the response deliberately.
Navigation after an action page.waitForNavigation() registered with the action Register both together to avoid missing a fast navigation.

For a click that triggers navigation, avoid starting the wait afterward. Register the wait and click together:

const [response] = await Promise.all([
  page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
  page.locator('a.next-page').click(),
]);

if (response && !response.ok()) {
  throw new Error(`Next page returned HTTP ${response.status()}`);
}

Puppeteer’s selector wait timeout defaults to 30 seconds unless changed. Set a timeout appropriate to the task and report which condition failed. Prefer waiting for a real element, response, or navigation state over inserting an arbitrary delay. A fixed delay can be too short on a slow run and waste time on a fast one.

5. Selectors, locators, and rendered content

The current Puppeteer interaction guide recommends locators for actions. Locators wait for the element to be present and for the state required by the action. CSS selectors work by default; Puppeteer also supports selector syntax for text, accessibility attributes, XPath, and Shadow DOM. See the page interactions guide for the current syntax and details.

For extraction, identify the element that represents the data you need and read its text or an attribute. Validate the result instead of assuming that a matching selector always contains useful data:

const cards = await page.locator('.product-card').mapAll(els =>
  els.map(el => ({
    name: el.querySelector('.product-name')?.textContent?.trim() ?? '',
    href: el.querySelector('a')?.href ?? '',
  }))
).wait();

const validCards = cards.filter(card => card.name && card.href);
console.log(validCards);

Use selectors tied to meaningful page structure where possible. A selector copied from a different page or a transient class name may stop matching after a site update. If expected content is missing, inspect the rendered page, frames, and shadow roots rather than treating an empty array as a successful scrape.

6. Handle responses and validate extracted data

page.goto() can return the main-resource response, which lets you check an HTTP status. A page can also render an error message or omit the expected content despite a successful navigation. Validate both the response where available and the extracted values.

const response = await page.goto('https://example.com/catalog', {
  waitUntil: 'domcontentloaded',
});

if (response && response.status() >= 400) {
  throw new Error(`HTTP ${response.status()} for ${page.url()}`);
}

await page.locator('.catalog-item').wait();
const count = await page.locator('.catalog-item').count();
if (count === 0) {
  throw new Error('Catalog loaded without any matching items');
}

For data loaded through a later request, wait for the specific response and check its status or body using the response API. Avoid logging credentials, session cookies, or sensitive page content when diagnosing failures.

7. Frames and Shadow DOM

Some pages put content in an iframe or a Shadow DOM tree. A selector evaluated against the main page may not reach content inside a separate frame. Inspect the page’s frames and use the relevant frame’s page APIs when the desired content lives there. For shadow-root content, use Puppeteer’s supported shadow DOM selector syntax or interact through a locator that targets it.

These are common reasons a selector can be valid yet return no result. Confirm where the content is rendered before broadening selectors or increasing timeouts.

8. Screenshots and PDFs

Puppeteer can capture a screenshot for visual debugging or produce a PDF from an HTML page. For example:

await page.screenshot({ path: 'page.png', fullPage: true });
await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });

page.pdf() uses print CSS by default, so the PDF can differ from the normal screen rendering. Creating a PDF from a page and navigating directly to an existing PDF document are different tasks; headless shell cannot navigate directly to a PDF document. Check the current Page API for screenshot and PDF options available in your installed Puppeteer version.

9. Reliability, performance, and cost

Keep browser work bounded

  • Close pages or the browser in a finally block so errors do not leave browser processes running.
  • Set explicit navigation and selector timeouts, and make failures identify the URL and stage that timed out.
  • Check the main response and validate the fields you extract. Record enough context to diagnose empty or malformed results.
  • For repeated work, plan browser lifecycle and concurrency deliberately. Each active page consumes resources; avoid opening unbounded pages at once.

Reduce unnecessary work

Do not wait for network idle if a specific result element is the actual readiness condition. Do not use a browser when the permitted source data is already available without rendering. Reuse a browser process carefully for a batch where appropriate, while ensuring each page is closed and state such as cookies does not leak between jobs.

Understand operating cost

Puppeteer itself is a library. Your practical costs come from the compute, browser installation, runtime, and maintenance required by the environment where it runs. The dossier provides no benchmark or universal resource estimate, so measure the pages and concurrency in your own deployment before sizing it.

10. Troubleshooting common failures

Symptom Likely cause Fix
Browser executable missing or launch fails Install scripts were blocked, or puppeteer-core was installed without a separately managed browser. Allow the required install step or run the documented browser install command; with core, configure a browser supplied by the environment.
Selector wait times out The page has not reached the expected state, the selector is wrong, or the content is inside a frame or shadow root. Inspect the rendered page and target structure; wait for the relevant state and query the correct frame or shadow content.
Scraper returns empty text The element matched before its content was populated, or the selector matched an empty/incorrect element. Wait for the actual content condition and validate the extracted value, not just selector presence.
Click succeeds but next page is missed The navigation wait was registered after the click and the navigation raced ahead. Start waitForNavigation() and the click together with Promise.all().
Navigation wait never completes The page does not navigate, or the chosen lifecycle event is not appropriate for the page. Use a selector or response wait if that is the real condition; select a suitable navigation lifecycle event and set a bounded timeout.
HTTP error despite a rendered page The main document response has an error status or the site rendered an error state. Inspect the response status and page content separately; do not treat rendering as proof of success.
Browser processes remain after an exception Cleanup was skipped on an error path. Put browser closure in a finally block.
PDF looks different from the browser view PDF rendering uses print CSS by default. Review print styles and PDF options; distinguish generating a page PDF from navigating to a PDF file.

11. Responsible access

Check the particular website’s published access rules and the requirements that apply to your use and location. Minimize the data you collect, keep it only as long as needed, and do not treat browser automation as authorization to access restricted material. The available research does not establish a universal legal rule, robots policy, or permission for any target site.

12. Or skip the browser setup

If your job is to capture a page rather than build and operate a browser scraper, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);
  • Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, and failed loads are never billed; response headers say which page verdict occurred and whether the request was billed. Cache hits cost nothing.
  • An MCP server lets AI agents, including Claude, Cursor, and other MCP clients, take screenshots.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Create a free ScreenshotNeo account and start with 1,000 screenshots a month, with no card required.

13. FAQ

Does Puppeteer scrape every website without JavaScript?

It can navigate ordinary HTML pages too, but a browser may add unnecessary work when the required content is already available without rendering or interaction.

Should I use puppeteer or puppeteer-core?

Choose puppeteer when you want its install to download a compatible Chrome. Choose puppeteer-core when your environment supplies and configures the browser.

Can Puppeteer download or parse a PDF?

page.pdf() creates a PDF from a page. That is distinct from navigating to an existing PDF document, which headless shell cannot do directly.