ScreenshotNeo

BlogHow-to

How to Scrape Custom Fields from JavaScript-Rendered SPAs

Extract custom fields from React, Vue, and Angular SPAs by capturing their APIs or waiting for the rendered DOM with a real browser.

By the ScreenshotNeo team29 September 202610 min read

How to Scrape Custom Fields from JavaScript-Rendered SPAs

Direct answer: a normal HTTP request often returns only a JavaScript application shell, so it cannot see fields that appear after the SPA runs. Use a real browser such as Playwright or Selenium, wait for the field’s populated state, and then extract it from either the API response or the rendered DOM. Prefer the API response when it contains the field: structured JSON is usually more stable than presentation selectors. Use the DOM when the value is computed in the browser or appears only after an interaction.

This guide shows a complete workflow for React, Vue, Angular, and similar SPAs, including authentication, interactions, pagination, retries, normalization, troubleshooting, and a managed alternative with ScreenshotNeo.

1. Decide where the custom field lives

Start by mapping one record from the page. Identify:

Choose the API response when it contains the field; otherwise read the rendered DOM after a semantic wait.
Choose the API response when it contains the field; otherwise read the rendered DOM after a semantic wait.
  • The route that displays the record or list.
  • The record identifier and its containing element.
  • The custom-field label or key.
  • The action that reveals the value: a tab, Details button, “load more” control, search, or scroll.
  • The network request that supplies the record.

Open browser developer tools and inspect the Network panel while performing the action manually. Look for JSON responses, GraphQL requests, or embedded data. If a response contains customField, meta, or a similar key, parse that response instead of scraping text. If the value is absent from all responses because it is calculated or formatted in the browser, read the rendered DOM.

2. Capture SPA API responses with Playwright

Playwright can observe requests and responses and wait for a response caused by navigation or an interaction. Register the listener before the action that triggers the request; otherwise a fast response can be missed. The official Playwright network documentation covers request and response events and page.waitForResponse().

Complete JavaScript example: extract JSON fields

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
  // Add storageState or extraHTTPHeaders here for authenticated apps.
});
const page = await context.newPage();

const responsePromise = page.waitForResponse(response =>
  response.url().includes('/api/records') &&
  response.request().method() === 'GET' &&
  response.status() === 200
);

await page.goto('https://example.com/records', {
  waitUntil: 'domcontentloaded'
});

const response = await responsePromise;
const payload = await response.json();

for (const record of payload.records ?? []) {
  console.log(JSON.stringify({
    id: record.id,
    customField: record.customField ?? null,
    sourceUrl: `https://example.com/records/${record.id}`
  }));
}

await browser.close();

Use a URL predicate that is specific enough to identify the intended request. For GraphQL, inspect response.request().postData() and match the operation name. For APIs that return several pages, capture the cursor or next link from each payload and continue until it is absent.

Trigger an interaction and capture its response

const detailsResponse = page.waitForResponse(response =>
  response.url().includes('/api/records/123/details') &&
  response.request().method() === 'GET'
);

await page.getByRole('button', { name: 'Details' }).click();
const details = await (await detailsResponse).json();
console.log(details.customField ?? null);

When the request is made by a service worker, page-level routing may not see it. Playwright documents that service-worker requests are not intercepted by page.route(). If interception is required, block service workers in the browser context or use context-level request handling after confirming that the application still functions.

3. Extract a field from the rendered DOM

Use the DOM when the custom field is not present in a useful response, is computed client-side, or is revealed only after several UI actions. Wait for evidence that the field is ready rather than sleeping for an arbitrary number of milliseconds.

Complete JavaScript example: scoped DOM extraction

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();

await page.goto('https://example.com/profile/123', {
  waitUntil: 'domcontentloaded'
});

const card = page.locator('[data-record-id="123"]');
await card.getByRole('button', { name: 'Details' }).click();

const field = card.locator('[data-field="customer-tier"]');
await field.waitFor({ state: 'visible' });

const value = (await field.textContent())?.trim() ?? null;
const href = await field.getAttribute('href');
console.log(JSON.stringify({ id: '123', value, href }));

await browser.close();

Prefer stable data-* attributes, accessible roles, labels, and names. Avoid generated CSS classes such as css-1a2b3c. Scope every selector to the record container so a duplicate label in a sidebar or another record cannot contaminate the result.

4. Authentication, cookies, and browser context

Authentication must exist in the same browser context that performs navigation. Choose the method used by the target application:

  • Login flow: navigate to the sign-in page, fill credentials from a secret store, submit, and wait for a post-login locator.
  • Saved session: create a Playwright storageState file in a controlled environment and load it into the context.
  • Headers: pass an authorization header when the application accepts one, while ensuring it is not logged.
  • Cookies: add required cookies to the context before opening the target route.
const context = await browser.newContext({
  storageState: 'playwright/.auth/user.json',
  extraHTTPHeaders: {
    Authorization: `Bearer ${process.env.API_TOKEN}`
  },
  timezoneId: 'UTC'
});

Do not place credentials in source code or captured output. Verify that you are allowed to access and collect the data, and check the site’s robots directives, terms, privacy requirements, rate limits, and applicable law.

5. Waiting correctly in React, Vue, and Angular

SPA rendering has several distinct readiness states. Pick the condition that proves the custom field is usable:

Situation Useful wait Why
Known API call page.waitForResponse() Confirms the data request completed.
Field has a stable element locator.waitFor({state:'visible'}) Confirms the value is rendered.
Loading indicator disappears locator.waitFor({state:'hidden'}) Useful when the app replaces a skeleton.
Several requests settle Wait for a specific response or a documented application marker Network idle alone can be misleading.

Avoid relying only on networkidle: analytics, polling, and advertisements can keep a page active forever, while cached data can make it appear idle before the field is populated. Set explicit timeouts and record which condition failed.

6. Scrolling, tabs, and lazy fields

A field may not exist until its panel opens or its row enters the viewport. Start the response promise, perform the action, then wait for the field:

const lazyResponse = page.waitForResponse(r =>
  r.url().includes('/api/records') && r.url().includes('cursor=')
);
await page.locator('[data-record-id="123"]').scrollIntoViewIfNeeded();
await page.getByRole('button', { name: 'Load more' }).click();
await lazyResponse;
await page.locator('[data-field="customer-tier"]').waitFor({ state: 'visible' });

For virtualized lists, scrolling may remove earlier rows from the DOM. Extract each row as it appears, or use the backing API and its cursor rather than assuming every record remains in the document.

7. Normalize and audit extracted values

Keep the difference between a missing field and an explicit null. Preserve the source URL, record ID, extraction timestamp, HTTP status, and page or cursor so a result can be replayed. Flatten nested objects only with a documented rule; otherwise retain the original JSON.

function normalize(record, sourceUrl, status) {
  return {
    id: String(record.id),
    customField: Object.hasOwn(record, 'customField')
      ? record.customField
      : undefined,
    sourceUrl,
    capturedAt: new Date().toISOString(),
    responseStatus: status
  };
}

Write failed record URLs and cursors to a durable queue. This lets you retry only failures instead of reprocessing an entire crawl.

8. Pagination, retries, and reliability

Follow the SPA’s own pagination mechanism: a cursor, next URL, page number, or “load more” request. Persist the cursor after each successful page. Use bounded retries with increasing delays for transient network errors and 5xx responses. Do not retry authentication failures or a consistent 4xx response without fixing the request.

  • Set a per-page timeout and an overall job deadline.
  • Limit concurrency to what the target and your host can support.
  • Deduplicate by stable record ID when retries overlap.
  • Log request URL, operation, status, duration, and failure reason without secrets.
  • Save HTML or JSON evidence only when permitted and protect personal data.

For a long-running process, pin browser and application dependencies, monitor browser crashes, and recreate a context after repeated failures. Keep a small replay set of known records to detect selector or schema changes.

9. Selenium alternative

Selenium is a good fit when your team already operates WebDriver or needs its browser-language ecosystem. Its official JavaScript quick start uses selenium-webdriver; Selenium Manager can handle browser-driver installation. Selenium also supports simulated user actions and arbitrary JavaScript execution.

import { Builder, By, until } from 'selenium-webdriver';

const driver = await new Builder().forBrowser('chrome').build();
try {
  await driver.get('https://example.com/profile/123');
  const button = await driver.findElement(By.css('[data-action="details"]'));
  await button.click();
  const field = await driver.wait(
    until.elementLocated(By.css('[data-field="customer-tier"]')),
    15000
  );
  console.log((await field.getText()).trim());
} finally {
  await driver.quit();
}

Compare Playwright and Selenium on browser coverage, network interception, locator quality, team language, hosting cost, and observability. Playwright generally provides a more direct response-waiting API; Selenium remains practical when WebDriver infrastructure is already standardized.

10. Common errors and fixes

Error Likely cause Fix
Empty HTML from HTTP client You received the application shell. Use a real browser or call the discovered JSON endpoint.
Response wait times out The listener was registered after navigation or the URL predicate is wrong. Create the promise before the action and inspect the actual request URL and method.
Field appears only after scrolling Lazy loading or virtualization. Scroll the row into view, then wait for the resulting response or locator.
Selector broke after redesign Generated class names changed. Use roles, labels, stable attributes, or the API schema.
Duplicate or stale value Selector matches another record or an old panel. Scope to the record ID and verify the payload’s ID.
Requests missing from interception A service worker made the request. Check service-worker behavior and use an appropriate context configuration.
Pagination gaps Cursor was not persisted or a page failed silently. Store every cursor and response status; replay failed pages.
CAPTCHA or bot-check page The site challenged automated browsing. Do not attempt to bypass it; follow the site’s access policy or obtain permission and an approved access method.

11. Performance and cost considerations

API extraction is usually cheaper and faster than opening a full browser page because it avoids layout, painting, and repeated interactions. Browser rendering is necessary when the API is unavailable, protected, or insufficient. Reuse a browser process and create contexts per job, limit parallel pages, and cache immutable responses. Select only the fields you need and stop waiting as soon as the field-specific condition succeeds.

Measure navigation time, response time, extraction time, browser memory, retry count, and records per page. A slower but deterministic workflow is easier to operate than one based on fixed sleeps and uncontrolled concurrency. Respect the target’s rate limits; a high request rate can cause throttling or block legitimate access.

12. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. It can capture a page after JavaScript runs, which is useful when you need visual evidence of a rendered SPA or want an AI agent to inspect the result. See the ScreenshotNeo API documentation for request options.

A managed browser can remove common overlays before producing a clean rendered capture.
A managed browser can remove common overlays before producing a clean rendered capture.
curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/records \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/records"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const fs = require('node:fs/promises');

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/records'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Before the capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

You can also use full-page capture with lazy images, CSS-selector element capture, custom JavaScript and CSS, clicks, selector or network-idle waits, custom headers and cookies, user agents, authorization, timezone and geolocation, request blocking, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and PDF output. The API accepts parameter names used by other screenshot services, which helps when switching.

Free accounts include 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

13. FAQ

Is scraping the rendered DOM always the best approach?

No. If the custom field is in a JSON or GraphQL response, parsing that response is less brittle and avoids presentation changes. Use the DOM for computed or interaction-only values.

Can I use a static HTTP client instead of a browser?

Only when the server response already contains the data. A static client does not execute the JavaScript that hydrates most SPAs.

How do I handle a field that is sometimes absent?

Represent missing and explicit null separately, and record the response schema or DOM state that led to the result. Do not convert every absence to an empty string.

What should I do when a site changes frequently?

Prefer API contracts, semantic locators, record-scoped selectors, replay fixtures, and alerts for extraction-rate changes. Keep failed URLs for quick diagnosis.

Can ScreenshotNeo return the custom field itself?

ScreenshotNeo returns screenshots, PDFs, and page information. Use it for rendered visual capture or agent inspection; use Playwright, Selenium, or the site’s permitted API when you need structured field values.