How to Scrape Public Data with Puppeteer
Learn how to collect specific fields from public pages with Puppeteer, using content-based waits, careful selectors, responsible access checks, and structured output.
How to Scrape Public Data with Puppeteer: launch a browser, navigate to a page you are permitted to access, wait for the specific content you need, extract only the required fields, validate the result, and close the browser. Puppeteer automates a real browser; it does not make public visibility or a successful request proof that collection is authorized. Check the site’s terms and applicable requirements, and do not bypass authentication, blocks, or other technical access controls.
This guide uses a small example that collects product names and prices from a page whose cards use .product-card, .product-name, and .price. Those selectors are illustrative: replace them with selectors from the public page you are allowed to process. The example returns structured JSON rather than saving full page markup.
1. Define the target and check access conditions
Before writing selectors, list the exact pages and fields required. Keep the collection scoped to that list. Review the target site’s terms and relevant legal requirements, especially if the data includes personal information or will be redistributed. Check its robots.txt instructions as part of crawler etiquette, but do not treat robots.txt as permission: RFC 9309 states that its rules “are not a form of access authorization” (RFC 9309).
If a page requires a login, presents a block, or otherwise restricts access, stop and resolve authorization through the site owner. Do not work around those controls. No target site is specified here, so its current terms, robots.txt rules, rate expectations, and allowed uses must be checked by you.
2. Install Puppeteer and create a script
Use a maintained Node.js release supported by the Puppeteer version you install. Puppeteer’s getting-started guide documents installing the package, launching a browser, creating a page, navigating, interacting, and reading page content (Puppeteer getting started). The package normally downloads a compatible Chrome for Testing browser; consult the installation guide if your environment manages browsers separately (Puppeteer installation).
mkdir public-page-scraper
cd public-page-scraper
npm init -y
npm install puppeteer
Save this as scrape.mjs. Set TARGET_URL to the public page you are authorized to access. The script waits for product cards, extracts only name and price, checks for empty values, emits JSON, and closes Chrome even if navigation or extraction fails.
import puppeteer from 'puppeteer';
const targetUrl = process.env.TARGET_URL;
if (!targetUrl) {
throw new Error('Set TARGET_URL to the public page you are authorized to access.');
}
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultNavigationTimeout(30_000);
page.setDefaultTimeout(10_000);
const response = await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
if (!response) {
throw new Error('Navigation produced no main-resource response.');
}
if (!response.ok()) {
throw new Error(`Page returned HTTP ${response.status()}`);
}
// Tie readiness to the content this task actually needs.
await page.waitForSelector('.product-card .product-name', { visible: true });
const products = await page.evaluate(() => {
return Array.from(document.querySelectorAll('.product-card')).map((card) => ({
name: card.querySelector('.product-name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null,
}));
});
if (products.length === 0) {
throw new Error('No product cards matched; inspect the page and selectors.');
}
const incomplete = products.filter((item) => !item.name || !item.price);
if (incomplete.length > 0) {
throw new Error(`${incomplete.length} product record(s) are missing a name or price.`);
}
process.stdout.write(`${JSON.stringify({ url: targetUrl, products }, null, 2)}\n`);
} finally {
await browser.close();
}
Run it with an environment variable. The URL must be quoted if it contains shell-special characters.
TARGET_URL='https://example.com/public-products' node scrape.mjs
The example checks the main navigation response and the extracted records, but a page can return HTTP 200 while showing an error, an empty state, or an unexpected layout. Inspect the actual result and add page-specific validation for the fields that matter.
3. Choose selectors that identify the intended data
CSS selectors are suitable when the page has stable classes or attributes. Keep a selector as specific as needed to avoid matching navigation, recommendations, or hidden templates. Puppeteer also supports custom selector syntax for text, accessibility attributes, XPath, and querying across open shadow roots (waitForSelector API; page interactions guide).
| Approach | Useful when | Check |
|---|---|---|
| CSS selector | The page has distinctive, reasonably stable classes or attributes. | Confirm it selects the intended element, ideally uniquely for a single field. |
| Text selector | A visible label or phrase identifies an element and is stable on that page. | Text may be localized, duplicated, or changed by editorial updates. |
| Accessibility selector | An accessible name or role provides a clear way to identify the target. | Validate that the name and role are present and unique on the target page. |
| XPath | A relationship between elements is easier to express as a document path. | Long paths tied to page structure can break when markup changes. |
Puppeteer recommends locators for selecting and interacting with elements; locators automatically wait for elements and relevant action conditions. Use a locator when the task involves an interaction such as clicking a visible control, and a content-specific wait before extraction (Puppeteer page interactions). For example:
await page.locator('button.load-more').click();
await page.waitForSelector('.product-card .product-name', { visible: true });
Only click controls that are part of the permitted, ordinary public-page flow. Avoid repeatedly clicking “load more” without a defined stopping condition.
4. Wait for the data, not an arbitrary delay
Choose readiness conditions based on how the target page renders:
page.goto(url, { waitUntil: 'domcontentloaded' })continues when the initial document has been parsed. It may be appropriate when the needed content appears shortly afterward through client-side code.page.waitForSelector(selector, { visible: true, timeout })waits for a matching visible element. It can return immediately if the selector is already present, and the timeout is configurable (waitForSelector API).page.waitForNetworkIdle({ idleTime })waits for network activity to become idle for at least the configured idle time. The consulted Puppeteer 25.12.0 API reference documents a 500 ms default idle time (waitForNetworkIdle API).page.waitForFunction(condition)can wait for a page-specific condition, such as a result count reaching the expected minimum. Make the condition bounded with an intentional timeout.
Network idleness is only a signal about network activity; it does not prove that the exact data has rendered. A page may keep analytics or long-polling connections open, or it may become idle before a delayed component appears. Prefer a selector or condition tied to the requested fields, and use network idle only when it matches the page’s behavior. API names and defaults can change across Puppeteer versions, so check the documentation for the version in your project.
5. Extract and validate structured fields
page.evaluate runs a function in the page context and returns its result; Puppeteer waits if that function returns a promise (Page.evaluate API). The function executes in the browser page, so pass any needed values as arguments rather than expecting it to access Node.js variables or imports.
const result = await page.evaluate((cardSelector) => {
return Array.from(document.querySelectorAll(cardSelector)).map((card) => ({
title: card.querySelector('h2')?.textContent?.trim() ?? null,
href: card.querySelector('a')?.href ?? null,
}));
}, '.article-card');
Keep the returned shape explicit. Normalize whitespace, represent absent values as null or flag them as invalid, and validate formats such as dates or prices in Node.js before saving. If data is repeated, decide whether duplicates are meaningful and define a stable key for de-duplication. Save only fields needed for the task and set an appropriate retention period.
6. Handle multiple pages carefully
For a small list of known public URLs, process each URL sequentially first. This keeps browser load and request frequency understandable. Check each page’s response and validate its extracted records independently so one malformed page does not silently contaminate the whole output.
const urls = [
'https://example.com/public-products?page=1',
'https://example.com/public-products?page=2',
];
const allResults = [];
for (const url of urls) {
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card .product-name', { visible: true });
const rows = await page.evaluate(() =>
Array.from(document.querySelectorAll('.product-card')).map((card) => ({
name: card.querySelector('.product-name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null,
}))
);
allResults.push({ url, rows });
}
Use the target site’s published pagination or navigation flow, and stop at a defined page or record boundary. Do not infer that a page count, request rate, or concurrency level is allowed just because the pages are publicly viewable. If you add retries, retry only failures plausibly caused by transient navigation or network problems; cap attempts, add backoff, and do not retry access denials or bot checks as a way to evade them.
7. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
TimeoutError waiting for a selector |
The selector is wrong, the component never rendered, it is hidden, or the page is slower than the configured timeout. | Inspect the page and selector; wait for a visible, content-specific element; adjust the timeout only when the page’s expected behavior justifies it. |
| Navigation has no response object | The navigation was interrupted or did not produce a main-resource response. | Log the URL and navigation error, inspect whether the page redirected or closed, and retry only if a transient failure is plausible. |
| HTTP error status | The server returned a non-success response, such as a missing page or a service restriction. | Check the URL and response status. Respect restrictions; do not attempt to defeat access controls. |
| Empty array despite a visible page | The data uses different markup, is inside a frame or shadow root, or has not rendered at extraction time. | Inspect the relevant DOM, verify selector uniqueness, and wait for a condition tied to the data. Use Puppeteer’s supported selector syntax where appropriate. |
| Records contain null fields | The page has optional fields, the selector matched a different card type, or the markup changed. | Decide whether the field is required; validate each record and update selectors to match the actual structure. |
| Network-idle wait never completes | Persistent connections or continuing background requests prevent the idle condition. | Use a selector or page-specific condition for the requested data rather than treating network idle as mandatory. |
| Browser fails to launch in a container | The browser binary or required system libraries are missing, or the runtime environment differs from local development. | Follow Puppeteer’s installation and deployment guidance for that environment; use a compatible browser installation and check launch diagnostics. |
| Page content looks like a block or challenge | The site is restricting automated access. | Stop collection and seek authorization or an official data interface. Do not bypass the challenge. |
8. Performance, reliability, and cost
Each browser process consumes memory and CPU, so avoid launching a new browser for every small page when a single controlled process can handle the job. Reuse a page for sequential navigations where appropriate, close pages and browsers in cleanup paths, and keep the set of requested fields small. For a larger workload, estimate resource use in your own permitted environment before raising concurrency; the research sources provide no benchmark or universal throughput figure.
Reliability comes from explicit waits, response checks, field validation, bounded retries, and recording enough context to diagnose a failed page. A timeout should not be converted into an empty successful record. Include the URL and failure category in logs, while avoiding sensitive headers or collected data that should not be retained. Since target layouts can change, periodically review extraction output for missing or malformed values.
Puppeteer itself is an open-source browser automation library, but running it has infrastructure costs: compute, memory, storage, and maintenance of the browser runtime. The exact cost depends on where and how often it runs; no general price or benchmark follows from the cited documentation. For one-off captures or workflows that need an image rather than extracted fields, an API can avoid operating browser infrastructure.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It is for capturing a page as an image or document; it does not replace Puppeteer when you need to extract arbitrary structured fields from a page. For screenshot work, cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 shots a month with no card, and paid plans start at $5 for 3,000 shots. Every feature is available on every plan. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
Replace the example URL with the page you are permitted to capture and use your API key. ScreenshotNeo’s parameter names used by other screenshot APIs also work, which can make switching easier. Create a free account for 1,000 screenshots a month with no card.
10. Frequently asked questions
Can Puppeteer scrape a page that requires JavaScript?
Yes. Puppeteer drives a browser and can read rendered page content after the relevant component appears. Wait for the content condition you need; simply waiting for the initial document does not guarantee client-rendered data is ready.
Is public data automatically okay to collect?
No. Public visibility does not settle authorization, terms, privacy, copyright, database rights, or other applicable requirements. Assess the specific site, data, use, and jurisdiction.
Should I use a fixed sleep?
Usually, use a selector or page-specific condition instead. A fixed delay can waste time on fast pages and still be too short on slow ones.
Can I use Puppeteer to get a screenshot instead of data?
Yes, Puppeteer can automate a browser for capture workflows. If you need a screenshot or PDF without setting up a browser, use the ScreenshotNeo API example above; if you need structured fields, keep the extraction and validation flow in Puppeteer.


