How to Scrape Dynamic Websites with JavaScript
Learn when to reproduce a page’s data request and when to use Playwright or Puppeteer, with runnable JavaScript examples, validation tips, and troubleshooting.
To scrape a dynamic website, first inspect the page’s network requests. If a repeatable request returns the data you need, reproduce that request in JavaScript; it is usually simpler than rendering the whole page. Use browser automation when the content depends on JavaScript execution, interaction, or the browser-rendered result. This guide uses Playwright for browser work and includes a request-based example, Puppeteer notes, validation, and operational guidance.
Scraping is not a way to bypass a site’s access controls. Before collecting data, check the site’s terms, access controls, privacy implications, applicable law, and intended use.
1. Inspect the page and choose an approach
Open the page in browser developer tools and inspect the Network panel while loading the page and performing the interaction that reveals the data. Look for a request whose response contains the fields you need. Check whether it is repeatable and whether it depends on session state or changing parameters.
| Approach | Use it when | Trade-off |
|---|---|---|
| Reproduce a data request | A repeatable request returns the desired data and can be understood and used appropriately. | Often transfers less data and requires less parsing than rendering a page. The request may rely on session state or change over time. |
| Browser automation | The page must execute JavaScript, requires interaction, or the result must match what a browser displays. | Provides browser behavior and DOM access, but requires browser setup and more resources. |
| Managed browser service | Operating browser instances or coordinating a site-wide crawl is a project requirement. | Moves some browser infrastructure work to a service. Check current capabilities and pricing in its documentation. |
Scrapy’s dynamic-content guide prefers reproducing requests containing the desired data when practical, because this can provide structured data with less parsing and network transfer. It identifies a headless browser as useful when reproducing the request is difficult or when the desired output exists only in the browser view. Scrapy: Selecting dynamically-loaded content.
2. Reproduce a data request with JavaScript
If inspection reveals a suitable JSON endpoint, request it directly. Adapt the URL and field names to the site you are authorized to access. This example uses Node.js built-in fetch (Node.js 18 or later):
const endpoint = new URL('https://example.com/api/products');
endpoint.searchParams.set('category', 'books');
const response = await fetch(endpoint, {
headers: { accept: 'application/json' },
signal: AbortSignal.timeout(15000),
});
if (!response.ok) {
throw new Error(`Request failed: HTTP ${response.status}`);
}
const payload = await response.json();
if (!Array.isArray(payload.items)) {
throw new Error('Unexpected response: items is not an array');
}
const rows = payload.items.map(item => ({
name: item.name ?? null,
price: item.price ?? null,
}));
console.log(rows);
For older Node.js versions or a project that already uses a request library, use that project’s HTTP client and apply the same checks: status, response format, expected fields, and a finite timeout. Do not assume that a request observed in developer tools is a stable public API. It may be an internal endpoint whose parameters or response change.
3. Scrape rendered content with Playwright
Use a browser when the needed content appears only after client-side execution or interaction. Install Playwright and its Chromium browser in a Node.js project:
npm install playwright
npx playwright install chromium
Save this as scrape.mjs. Replace the example URL and selector with values observed on the target page. The locator wait is tied to evidence that the target content appeared rather than an arbitrary sleep.
import { chromium } from 'playwright';
const url = 'https://example.com/catalog';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
page.setDefaultTimeout(10000);
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30000,
});
if (!response) throw new Error('Navigation returned no response');
if (!response.ok()) {
throw new Error(`Navigation failed: HTTP ${response.status()}`);
}
const cards = page.locator('[data-testid="product-card"]');
await cards.first().waitFor({ state: 'visible' });
const products = await cards.evaluateAll(elements =>
elements.map(element => ({
name: element.querySelector('h2')?.textContent?.trim() ?? null,
price: element.querySelector('[data-price]')?.textContent?.trim() ?? null,
href: element.querySelector('a')?.href ?? null,
}))
);
if (products.length === 0) {
throw new Error('No products found; check the selector and page state');
}
console.log(JSON.stringify(products, null, 2));
} finally {
await browser.close();
}
Install the browser binary in the environment where the script runs, not just on a developer laptop. For production, pin the Playwright dependency and deploy a compatible browser version. The Playwright Page API documents navigation, page events, request observation and routing, and waits.
Wait for the right condition
- Wait for a locator: use
locator.waitForor a locator action when an element’s presence or action readiness demonstrates that the page is ready. Playwright’s locator APIs perform auto-waiting for actions. - Wait for a known response: if a specific request supplies the data, observe or wait for that response, then validate its body. This can be more direct than waiting for unrelated page activity to stop.
- Choose navigation readiness deliberately:
domcontentloadedmeans the initial document has been parsed, not that a client-rendered list has appeared. Wait separately for the content you need. - Avoid fixed sleeps as the readiness test: a short pause can be too short on a slow response and unnecessarily long on a fast one.
Observe requests when the data source is unclear
Playwright can observe network requests and responses. Use that capability to identify what supplies the content, then decide whether a direct request is appropriate or whether the browser must remain in the workflow. Route requests only when the change is part of your intended, authorized behavior; blocking resources can also prevent the page from working.
4. Puppeteer alternative
Puppeteer is another JavaScript browser automation option. Its guide recommends locator-based interaction because locators wait for the element to be present and ready for the action. Use the same workflow: navigate, wait for a meaningful locator, extract only the fields needed, and close the browser in a finally block. See the Puppeteer page interaction guide.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/catalog', {
waitUntil: 'domcontentloaded',
timeout: 30000,
});
const cards = await page.locator('[data-testid="product-card"]').wait();
const products = await page.$$eval('[data-testid="product-card"]', elements =>
elements.map(element => ({
name: element.querySelector('h2')?.textContent?.trim() ?? null,
price: element.querySelector('[data-price]')?.textContent?.trim() ?? null,
}))
);
console.log(products);
} finally {
await browser.close();
}
The locator wait above establishes that at least one matching element is present before extraction. Adapt it if the page legitimately has no results; a zero-result state should be detected and represented explicitly rather than treated as successful extraction.
5. Extract, validate, and store results
- Extract narrowly. Select the fields the job needs instead of serializing the entire DOM.
- Handle missing values. A selector can disappear or a field can be absent. Return
nullor record a row-level error instead of crashing the whole batch without context. - Validate a sample. Compare a few extracted records with the rendered page or source response. Confirm that prices, dates, links, and pagination mean what your dataset expects.
- Record provenance. Store the source page and retrieval time with the data so a later change can be traced.
- Separate extraction from persistence. Write validated records through a controlled storage layer, and avoid losing an entire run because a single row is malformed.
These validation and schema practices are implementation guidance; there is no universal extraction schema. Treat page markup and observed endpoints as changeable inputs, not contracts unless the site documents them as such.
6. Pagination, interaction, and crawl scope
For a page with a next button, wait for the current results to change after clicking and set a page or record limit. For infinite scrolling, scroll in bounded increments and stop when the item count no longer grows or the intended scope is reached. Deduplicate by a stable record identifier or canonical URL when possible. Do not let a pagination loop run without a maximum page count, time limit, and failure policy.
For multiple pages, begin with a small sample and measure the actual time and failure modes in your environment before choosing concurrency. Keep concurrency bounded, use backoff for transient failures, and respect the site’s policies and access controls. A site-wide crawl may call for a crawl-oriented system rather than maintaining a browser session per page. Cloudflare documents Browser Run Quick Actions, browser sessions controllable through Playwright, Puppeteer, CDP, or Stagehand, and a crawl endpoint for site-wide extraction. Check its current Browser Run documentation for availability and plan details.
7. Reliability, performance, and cost
- Prefer the smallest suitable operation. A direct structured request can avoid downloading and rendering assets. Browser automation is appropriate when page execution or interaction is necessary.
- Wait on page evidence. Condition-based waits reduce needless delays and make failures more diagnosable than a single long fixed pause.
- Set timeouts at each boundary. Navigation, element waits, and HTTP requests can each stall. Give each a finite limit and log which stage timed out.
- Bound resource use. Browser processes consume more resources than a simple HTTP request. Reuse a browser for a controlled batch where appropriate, isolate page state, and always close pages and browsers after errors.
- Limit concurrency and retries. Excessive parallelism can burden the target and your own workers. Retry only transient failures with a cap and backoff; do not repeatedly retry a persistent selector or permission error.
- Track changes. Log the URL, run time, response status, extracted record count, and a concise failure reason. Alert on sudden empty results or schema changes.
- Budget operational cost from actual workload. Account for browser compute, network transfer, storage, retries, and any managed service charges. The sources here establish capabilities, not comparative benchmarks, so no universal speed or cost figure applies.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Selector wait times out | Wrong selector, content has not loaded, or the page is in a different state. | Inspect the live DOM and request sequence; wait for the actual result condition and handle empty states separately. |
| Navigation succeeds but extracted list is empty | domcontentloaded occurred before client rendering, or the site returned a different page. |
Wait for a result locator or relevant response. Check status, final URL, and a small amount of page text. |
| Works locally but browser launch fails in deployment | The compatible browser binary or system dependencies are absent, or the deployed version differs. | Install the browser for the pinned automation package in the deployment image and keep package/browser versions aligned. |
| HTTP 401 or 403 | The request lacks authorized session context or the site denies access. | Confirm that access is permitted and use only legitimate credentials and documented flows. Do not attempt to defeat access controls. |
| HTTP 429 | Requests are being rate-limited. | Reduce request rate and concurrency, honor any published guidance, and apply capped backoff. Stop if access is not allowed. |
| Unexpected JSON parse error | The endpoint returned HTML, an error body, or a response format that changed. | Check status and content type before parsing, capture a safe diagnostic sample, and update validation for the observed schema. |
| Duplicate or missing records across pages | Pagination state was not synchronized, pages overlap, or scrolling stopped too early. | Wait for a page-specific change, deduplicate by a stable key, and set explicit page and item limits. |
| Run hangs or uses excessive memory | Unbounded navigation, pages not closed, excessive concurrency, or too much DOM captured. | Add per-stage timeouts, close resources in cleanup paths, bound concurrency, and extract only required fields. |
9. Responsible crawling
Google documents that its automated crawlers use the Robots Exclusion Protocol and that robots.txt rules apply to the host, protocol, and port of the robots.txt file. This describes Google’s crawler behavior and the scope of those rules; it does not settle every scraper’s legal, contractual, or privacy obligations. Check the target site’s own terms, access controls, privacy implications, applicable law, and your intended use before collecting data. See Google’s robots.txt introduction.
10. Or skip the browser setup
If your goal is a clean rendered screenshot rather than extracting a dataset, ScreenshotNeo provides a website screenshot API and MCP server for developers. A single GET request can return a PNG, JPEG, WebP, or PDF. The JavaScript examples above show how to scrape page data; this API captures the page as an image or PDF.
See the ScreenshotNeo API documentation for parameters and setup. Example calls for the same page:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan.
Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card required.
FAQ
Can JavaScript scrape a page without a browser?
Yes, when the desired data is returned by a request you can reproduce. If the content requires client-side execution or interaction, use browser automation or another suitable rendering workflow.
Should I use Playwright or Puppeteer?
Both provide browser automation. Choose based on the APIs and runtime already used by your project, then use locator waits, bounded timeouts, and explicit validation.
Does robots.txt grant permission to scrape?
No. It is a crawler convention with defined scope; it does not answer all legal, contractual, privacy, or access questions.
Can ScreenshotNeo extract structured product data?
The documented ScreenshotNeo offering described here returns screenshots or PDFs. Use a data request or browser DOM extraction when you need structured records.


