How to Use Playwright for Web Scraping
Learn when Playwright fits a scraping job, how to wait for dynamic content, extract and validate records, and troubleshoot common failures.
Use Playwright for web scraping when a page needs browser rendering or interaction before its content is available. Navigate to the page, wait for a meaningful signal, locate the fields you need, extract them, and validate the results before saving. If the response already contains the content, a regular HTTP request and HTML parser may be simpler.
This guide uses Node.js for the runnable Playwright examples. It covers browser setup, dynamic pages, pagination, structured extraction, diagnostics, and common failures. Collect only data you are permitted to access; Playwright’s technical documentation does not grant permission to scrape a particular site.
1. Decide whether Playwright is the right tool
Start by checking whether the page’s initial HTTP response contains the information you need. If it does, an HTTP client and HTML parser can avoid launching a browser. Use Playwright when the page renders the data in the browser, requires a user interaction to reveal it, or depends on browser behavior.
| Approach | Use it when | Tradeoff |
|---|---|---|
| HTTP request and parser | The response HTML contains the needed fields. | Less browser setup; it cannot perform browser interactions or render client-side content by itself. |
| Playwright | Content appears after rendering or interaction, or the task depends on browser behavior. | Requires browser automation and its associated setup and resource use. |
This is a design comparison, not a benchmark. The reviewed sources provide no measured speed or cost comparison. Before collecting data, check the target site’s terms, access controls, and applicable requirements. An endpoint observed in browser traffic is not automatically permission to reuse its data.
2. Install Playwright and run a first scraper
Create a Node.js project, install Playwright, and install its Chromium browser. These commands use the Playwright package’s documented import style:
npm init -y
npm install playwright
npx playwright install chromium
Save this as scrape.mjs. Replace the example URL and selectors with ones from a site you are allowed to access. The [data-product-name] attribute is an illustrative selector, not a claim that it exists on the example page.
import { chromium } from 'playwright';
const browser = await chromium.launch();
try {
const page = await browser.newPage();
const response = await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
const heading = await page.getByRole('heading', { name: 'Catalog' }).textContent();
const names = await page.locator('[data-product-name]').allTextContents();
const records = names.map(name => ({ name: name.trim() })).filter(row => row.name);
if (records.length === 0) {
throw new Error('No product names found; check the selector and page state.');
}
console.log(JSON.stringify({ heading: heading?.trim() ?? null, records }, null, 2));
} finally {
await browser.close();
}
Run it with node scrape.mjs. Closing the browser in finally ensures it is closed when navigation or extraction throws an error. For one-off scripts, this is a useful baseline; a scheduled or long-running scraper should also record failures and avoid saving incomplete records.
3. Choose locators that survive page changes
Playwright recommends locators based on how users identify content and controls, such as roles, labels, and visible text. Deep CSS or XPath chains tied to layout can break when the page structure changes. A stable, explicit data attribute can also be appropriate when the site provides it as a reliable contract.
// Prefer a user-facing locator when the page exposes the role and name.
const title = await page.getByRole('heading', { name: 'Catalog' }).textContent();
const next = page.getByRole('link', { name: 'Next page' });
// A data attribute can be useful when it is present and stable.
const productNames = await page.locator('[data-product-name]').allTextContents();
Inspect the actual page to determine its roles, accessible names, and attributes. Placeholder selectors and names in examples must be adapted to the target.
4. Wait for content, not an arbitrary delay
Navigation completing does not necessarily mean a client-rendered list is ready. Wait for a concrete page state, such as a result list becoming visible. Playwright locators auto-wait and retry for actions; its guidance favors locator waits or web assertions over the older waitForSelector approach.
await page.goto('https://example.com/catalog');
await page.getByRole('list', { name: 'Products' }).waitFor({ state: 'visible' });
const rows = await page.getByRole('listitem').allTextContents();
Use the real role and accessible name exposed by the page. If a known loading indicator disappears before results are ready, that can be a useful signal, but confirm the results themselves are present too. Fixed sleeps can finish too early on a slow response and waste time on a fast one; use them only when the site provides no better readiness signal.
Choose navigation timing deliberately. domcontentloaded waits for initial document parsing and is often a reasonable point to begin waiting for a specific element. A page that needs a particular request or client-side render still requires its own readiness condition. Avoid assuming that one navigation event means all page data is ready.
5. Extract records and validate them
Keep the output schema small and explicit. For example, a catalog record might need a name, price, and canonical URL. Check required values and plausible formats before writing a record; treat missing values as a signal to inspect the page rather than silently saving bad data.
const cards = page.locator('[data-product-card]');
const count = await cards.count();
const records = [];
for (let i = 0; i < count; i++) {
const card = cards.nth(i);
const name = (await card.locator('[data-product-name]').textContent())?.trim();
const price = (await card.locator('[data-price]').textContent())?.trim();
const href = await card.locator('a').first().getAttribute('href');
if (!name || !href) continue;
records.push({ name, price: price || null, href: new URL(href, page.url()).href });
}
if (records.length === 0) {
throw new Error('No valid records extracted.');
}
console.log(records);
The selectors above are examples; verify they exist on the actual page. For traceability, store the source page URL and retrieval time alongside each batch. If required fields are absent, the record count is unexpectedly low, or duplicates appear, flag the run for review instead of treating the extraction as complete.
6. Handle pagination and interaction
When a page exposes a next-page control, use its locator and wait for the next result state after clicking. Add a page limit or another stopping condition so a changed or misidentified control cannot create an unbounded loop.
const maxPages = 20;
const seen = new Set();
const allNames = [];
for (let pageNumber = 0; pageNumber < maxPages; pageNumber++) {
await page.getByRole('list', { name: 'Products' }).waitFor({ state: 'visible' });
const names = await page.locator('[data-product-name]').allTextContents();
for (const raw of names) {
const name = raw.trim();
if (name && !seen.has(name)) {
seen.add(name);
allNames.push(name);
}
}
const next = page.getByRole('link', { name: 'Next page' });
if (await next.count() === 0 || !(await next.isVisible())) break;
await next.click();
}
console.log(allNames);
Adapt the stop condition to the page: a disabled next button, a changed URL, or a known final-page marker may be more reliable than checking whether a link exists. If clicking does not change the content, wait for a specific page change or inspect whether the control is disabled or intercepted.
7. Inspect network activity when rendering is unclear
Playwright can observe and route HTTP and HTTPS traffic, including XHR and fetch requests. Network events can help diagnose when a page obtains data or help test an application you own. They do not establish that collecting or reusing an observed endpoint is allowed; check the target site’s terms and access requirements.
page.on('response', response => {
const request = response.request();
if (request.resourceType() === 'xhr' || request.resourceType() === 'fetch') {
console.log(response.status(), response.url());
}
});
await page.goto('https://example.com/catalog');
Register listeners before navigation so early requests are visible. Keep diagnostic logging scoped: URLs and response details may contain sensitive information. For multiple pages that share settings, a BrowserContext provides a place for shared browser configuration and can host multiple pages.
8. Configuration and operational choices
- Browser: launch the browser engine your task needs; install that browser in the environment where the scraper runs.
- Page and context: use a context when multiple pages should share settings, and a separate context when you need isolation between sessions.
- Navigation timeout: set a bounded timeout appropriate to the job, then report a timeout as a failed or incomplete retrieval rather than as an empty successful result.
- Readiness: wait for a meaningful locator or expected page state after navigation and interaction.
- Concurrency: begin with a small number of pages and increase only when the target and your runtime can handle it. No benchmark in the reviewed sources establishes a universally safe or faster concurrency value.
- Output: validate required fields, retain source URLs and retrieval timestamps, and make retries safe so a repeated run does not create accidental duplicates.
For reliability, distinguish navigation failures, access-denied or challenge pages, empty results, and valid records in logs. Do not turn a failed page into a successful empty dataset. For performance and cost, account for browser startup and resource use as operational tradeoffs; there are no sourced benchmark figures here. If a static response is sufficient, a plain request and parser avoids browser setup.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chromium executable is missing | The Playwright package is installed, but its browser was not installed in this environment. | Run npx playwright install chromium in the deployment environment. |
| Navigation times out | The page is slow, unreachable, or waiting for a stronger load condition than the task needs. | Use a bounded timeout, choose an appropriate navigation event, and wait separately for the actual content locator. Record the timeout as a failed retrieval. |
| Extraction returns no text | The selector is wrong, content has not rendered, or the page structure changed. | Inspect the rendered page, verify the locator, and wait for a concrete result state before extraction. |
| Only some records appear | Content loads incrementally, pagination was missed, or required fields are absent. | Wait for the result state, handle pagination with a stopping condition, and validate the required fields and expected count. |
| Click does not advance results | The locator may match the wrong control, the control may be disabled, or the page may need a different state transition. | Check visibility and enabled state, use a precise accessible locator, and wait for a changed URL or result state after the click. |
| Output is empty but the script succeeds | The script treats an access-denied page or failed response as a valid empty result. | Check navigation response and visible page state; fail or flag the run when required content is absent. |
| Selector breaks after a redesign | The selector is coupled to internal layout structure. | Prefer role, label, or text locators, or a stable explicit data attribute, and update the extraction schema when the page changes. |
10. Or skip the browser setup
If your goal is a screenshot rather than extracting structured records, ScreenshotNeo takes a website URL and returns an image or PDF. Its API accepts a GET request; see the ScreenshotNeo API documentation for the available options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
11. Frequently asked questions
Can Playwright scrape a page that requires sign-in?
Playwright can automate browser behavior, but whether you may access or collect the page’s data depends on the site’s rules and applicable requirements. This guide does not determine permission for any specific site.
Should I use a fixed delay after every page load?
Usually, wait for a locator or meaningful state instead. A fixed delay can be too short on a slow page and unnecessarily long on a fast one.
Does seeing a data request in DevTools or Playwright make it okay to call directly?
No. Observing network traffic is a diagnostic capability, not authorization to collect or reuse data from that endpoint.
When is a screenshot API a better fit?
When the deliverable is a screenshot or PDF rather than structured records, a screenshot API can avoid maintaining browser setup in your own scraper.
Sources
- Playwright locators for locator guidance and auto-waiting.
- Playwright Page API for navigation and page operations.
- Playwright network documentation for observing and routing HTTP and HTTPS traffic.
- Playwright browser contexts for context and page organization.


