Puppeteer Web Scraping: A Practical Guide
Learn when Puppeteer is the right scraper, how to wait for rendered content, extract reliable results, handle failures, and stay within access rules.
Puppeteer is a good choice when the data you need appears only after a browser runs JavaScript or interacts with a page. Launch a browser, navigate to the page, wait for a specific signal that the relevant content is ready, extract only the fields you need, validate the result, and close the browser. If the required HTML or JSON is already available from a direct HTTP response, parsing that response is usually simpler.
Puppeteer controls Chrome or Firefox through supported automation interfaces and runs headless by default. It is a browser automation library, not permission to collect data from any particular site. Review the site’s terms and applicable rules before collecting data.
1. Decide whether you need a browser
Start by checking the response you can obtain without rendering a page. If it contains the records you need, a direct HTTP request and an HTML or JSON parser avoid the browser setup. Use Puppeteer when browser-side JavaScript, a user interaction, or a browser-specific state is necessary to expose the data.
Use the narrowest method that meets the requirement. A page that loads data through an API may be easier to understand by inspecting the page’s network activity and then using an authorized, documented endpoint where available. Do not assume an endpoint is public or permitted simply because the browser can request it.
2. Install Puppeteer and prepare the browser
The normal puppeteer installation path downloads a compatible browser. puppeteer-core installs the library without downloading a browser; choose it when you manage the browser separately or connect to an existing browser. If your package manager or deployment environment blocks install scripts, follow Puppeteer’s browser installation instructions and make sure the browser is present at runtime.
npm init -y
npm install puppeteer
For an ES module script, set the package type in package.json:
{
"type": "module",
"scripts": {
"scrape": "node scrape.js"
}
}
Save the following as scrape.js. Replace the URL and placeholder selectors after inspecting a page you are authorized to access. This example waits for the content signal, extracts a few fields, checks the result, and closes the browser even if an operation fails.
import puppeteer from 'puppeteer';
const url = 'https://example.com/catalog';
const cardSelector = '.product-card';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
page.setDefaultTimeout(15_000);
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30_000
});
if (response && !response.ok()) {
throw new Error(`Page returned HTTP ${response.status()}`);
}
// Wait for the page-specific signal that records are ready.
await page.locator(cardSelector).wait();
const records = await page.$$eval(cardSelector, cards =>
cards.map(card => ({
title: card.querySelector('.title')?.textContent?.trim() ?? '',
url: card.querySelector('a')?.href ?? ''
}))
);
if (records.length === 0) {
throw new Error(`No records found for selector ${cardSelector}`);
}
for (const record of records) {
if (!record.title || !record.url) {
throw new Error(`A record is missing a title or URL: ${JSON.stringify(record)}`);
}
}
console.log(JSON.stringify(records, null, 2));
} finally {
await browser.close();
}
Run it with npm run scrape. The selectors are examples, not universal selectors: inspect the target page and choose elements that correspond to the data you need.
3. Wait for the state your task needs
A navigation event or quiet network does not prove that the records you need are present. Choose a wait condition tied to the actual task, then independently validate the resulting URL, response, or DOM state.
| What must be ready | Useful signal | What to verify afterward |
|---|---|---|
| A content element | page.locator(selector).wait() or waitForSelector |
Expected count and required fields |
| A custom DOM condition | waitForFunction |
The condition still holds when extracting |
| A document or URL transition | waitForNavigation, registered before the action |
Expected URL, response status, and page content |
| A particular server response | waitForResponse with a narrow predicate |
Expected UI state and data, not just the request |
| An iframe | Wait for the frame to appear | Query the intended frame and verify its contents |
| Late page resources | waitForNetworkIdle |
Confirm the target data is present and correct |
For an explicit result-count condition, use waitForFunction:
const minimumResults = 10;
await page.waitForFunction(
count => document.querySelectorAll('.result-row').length >= count,
{ timeout: 15_000 },
minimumResults
);
For a page that updates after a server response, wait for a narrowly identified response and then confirm the DOM. A response wait proves a matching request completed; it does not by itself prove the page rendered the expected result.
const responsePromise = page.waitForResponse(response => {
const request = response.request();
return response.url().includes('/api/catalog') &&
request.method() === 'GET';
}, { timeout: 15_000 });
await page.locator('button.refresh').click();
const apiResponse = await responsePromise;
if (!apiResponse.ok()) {
throw new Error(`Catalog request returned HTTP ${apiResponse.status()}`);
}
await page.locator('.result-row').wait();
Make the URL or endpoint predicate as specific as the task allows. A broad predicate may be satisfied by unrelated page traffic.
4. Click through navigation safely
When a click may navigate to another document or change a single-page application’s URL, start waiting for the transition before clicking. Then check the result. Same-document transitions can produce a null response, so validate the URL or expected page content too.
const [response] = await Promise.all([
page.waitForNavigation({ waitUntil: 'domcontentloaded' }),
page.locator('a.next-page').click()
]);
if (response && !response.ok()) {
throw new Error(`Unexpected status: ${response.status()}`);
}
if (!page.url().includes('page=2')) {
await page.locator('.page-two-results').wait();
}
Use this pattern only when the action can cause a navigation. For an in-place update, wait for the relevant response or DOM change instead. For pagination, keep a finite page limit or another clear stopping condition, and detect repeated URLs or repeated records so a changed page does not cause an infinite loop.
5. Select and extract only the fields you need
CSS selectors are the usual starting point. Puppeteer also supports custom selector syntax for XPath, text, accessibility attributes, and Shadow DOM. Use locators for ordinary interactions; they wait and retry around action readiness conditions such as visibility, enabled state, viewport placement, and a stable bounding box. Use $eval and $$eval when the relevant elements already exist and you want to read values from them.
const titles = await page.$$eval('.product-card .title', nodes =>
nodes.map(node => node.textContent?.trim() ?? '')
);
const firstTitle = await page.$eval(
'.product-card .title',
node => node.textContent?.trim() ?? ''
);
Prefer selectors tied to meaningful attributes or content where possible. A selector is an assumption about a page’s structure, and a site can change it. Validate extracted values, record empty or malformed fields, and review selectors when results unexpectedly change.
If you use waitForSelector and keep its returned element handle, dispose of the handle when finished. After a full document replacement, query the new document rather than reusing handles from the old one.
6. Handle iframes and dynamic content
Content inside an iframe belongs to a separate frame. Wait until the frame is present, identify it reliably, and then query that frame. Do not assume a selector on the main page can reach into an iframe.
await page.waitForFunction(() =>
[...document.querySelectorAll('iframe')].some(frame =>
frame.title === 'Catalog results'
)
);
const frame = page.frames().find(candidate =>
candidate.url().includes('/embedded-catalog')
);
if (!frame) throw new Error('Catalog iframe did not appear');
await frame.locator('.result-row').wait();
const rows = await frame.$$eval('.result-row', elements =>
elements.map(element => element.textContent?.trim() ?? '')
);
Cross-origin frames can still be automated as frames by the browser, but their contents are not part of the parent document. Locate the correct frame and handle its own loading and errors. If a site replaces an iframe during an update, reacquire the frame instead of relying on a stale reference.
7. Validate collection and bound the work
Extraction should fail visibly when it returns an unexpected state. At minimum, validate that results are nonempty where expected, fields have the right shape, and the current URL or page identity matches the intended record set. For paginated collection:
- Set a maximum number of pages or a maximum total record count.
- Stop on a clear end-of-results signal.
- Track visited page URLs to detect loops.
- Deduplicate records using a stable identifier when one is available.
- Log the page URL and the reason collection stopped.
Keep a distinction between an empty but valid result and a failed page load. A challenge page, an access-denied page, a changed selector, and a genuinely empty catalog can all produce zero records; inspect the response and page state rather than treating every empty array the same.
8. Request interception: use it carefully
Request interception can block or modify requests, but once enabled, every intercepted request must be continued, responded to, aborted, or served from cache. A handler that leaves even one request unresolved can stall page activity and make waits time out.
await page.setRequestInterception(true);
page.on('request', request => {
if (request.resourceType() === 'image') {
void request.abort().catch(error => {
console.error('Could not abort image request:', error.message);
});
return;
}
void request.continue().catch(error => {
console.error('Could not continue request:', error.message);
});
});
Use interception only when the task benefits from controlling requests, and define how every request is handled. Blocking resource types can also remove content or alter page behavior, so verify that the fields you extract remain available. Avoid adding interception as a generic speed trick without measuring the actual task.
9. Responsible access and reliability
RFC 9309 defines the Robots Exclusion Protocol. It describes crawler rules that site owners request crawlers to honor and states: “These rules are not a form of access authorization.” Read and honor relevant robots.txt guidance, but do not treat it as permission to access data or as a substitute for the site’s terms, authorization, or applicable legal review. RFC 9309.
The US Supreme Court’s decision in Van Buren v. United States interpreted “exceeds authorized access” in the context of a law-enforcement database and access to areas off limits to the user. It did not decide that scraping any public website is lawful. The decision is specific to its legal context; terms, technical restrictions, privacy and intellectual-property rules, the data and purpose, and jurisdiction can all matter. For consequential projects, review the applicable terms and get qualified legal advice. Supreme Court opinion in Van Buren v. United States.
For reliable runs, distinguish timeouts, non-OK responses, unexpected URLs, empty results, and extraction failures in logs. Close the browser in a finally block, keep collection bounded, and avoid retrying indefinitely. A retry can help with transient failures, but it cannot fix a changed selector, a blocked request, or a page that no longer exposes the expected data.
10. Performance, reliability, and cost
A browser has more setup and runtime work than parsing an already available HTTP response. Puppeteer’s practical cost depends on the browser, pages, navigation, waits, and work performed; the reviewed sources do not provide a universal speed or cost benchmark. Keep the job bounded, reuse a browser for a controlled batch when appropriate, and close pages and browser processes when work ends. Measure your own authorized workload rather than assuming a fixed performance gain.
Make waits task-specific and give them finite timeouts. A fixed delay can waste time when the page is ready early and still fail when it is late. Network-idle waits can be unsuitable for pages with persistent requests, and network quiet alone does not establish data correctness. Record enough context to reproduce a failure: target URL, wait condition, status, selector, and whether the expected content appeared.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser executable not found | Install scripts were blocked, or a separately managed browser is missing. | Install the compatible browser using Puppeteer’s official setup instructions; verify the runtime can find its executable. If using puppeteer-core, provide or configure the browser you manage. |
| Navigation or selector timeout | The page did not reach the assumed state, the selector changed, or the timeout is too short. | Inspect the final URL and page state; wait for a task-specific element, response, or condition; confirm the selector against the current page. |
| Navigation wait hangs after a click | The click updated the page in place rather than navigating, or the wait was registered too late. | Register navigation waits before the click. For same-document or in-place changes, wait for a URL, response, or DOM condition instead. |
| Wait resolves but extracted data is empty | The signal was too broad, content is in an iframe, or the selector targets the wrong structure. | Use a narrower readiness condition, inspect frames, and validate counts and field values before accepting the result. |
| Unexpected status or access-denied content | The server returned an error or a different page than expected. | Check the response status and resulting URL, respect the site’s access rules, and do not treat browser automation as a way around restrictions. |
| Requests stall after enabling interception | A request event was not handled, or a handler failed. | Ensure every request is continued, responded to, aborted, or served by cache; log handler errors and test that page resources still load. |
| Fields become undefined after navigation | Code is using an element handle from a replaced document. | Query the current document again and dispose of handles that are no longer needed. |
| Pagination repeats or runs indefinitely | The stop condition is missing or the page transition did not advance. | Set a maximum page count, track visited URLs, verify page identity after each transition, and stop on a clear end signal. |
For Puppeteer’s current installation, locator, selector, and wait APIs, use the official Puppeteer documentation.
Or skip the browser setup
If your task is to capture a page rather than extract structured records, ScreenshotNeo can return an image or PDF with one GET request. Its API supports PNG, JPEG, and WebP screenshots and PDFs, and its documentation covers request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
With Node.js on a runtime that does not provide Bun.write, save the response body using that runtime’s file API. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots; all features are on every plan. See ScreenshotNeo for plan details and the API documentation for options.
Create a free account for 1,000 screenshots a month, with no card.
FAQ
Does Puppeteer scrape data automatically?
No. It automates a browser; you still need to identify the page state, selectors, fields, and stopping conditions for the task.
Is a successful navigation proof that data is ready?
No. Check the response and URL where relevant, then verify the expected DOM state or data before extraction.
Can robots.txt authorize scraping?
No. RFC 9309 explicitly says robots.txt rules are not access authorization. Consider site terms, authorization, and applicable laws separately.
Should I use Puppeteer for a screenshot?
Use it when you need browser automation as part of a larger workflow. For a screenshot or PDF without managing a browser, ScreenshotNeo offers a one-call API and an MCP server for AI agents.


