ScreenshotNeo

BlogHow-to

Web Scraping Single-Page Applications with Python and Headless Browsers

Learn to scrape JavaScript-heavy SPAs with Python, Playwright, network inspection, direct API calls, retries, and compliance checks.

By the ScreenshotNeo team1 October 202610 min read

Short answer: use a real browser such as Playwright when the page must execute JavaScript, then wait for application state rather than the initial HTML. Inspect the XHR and fetch requests that contain the data. If a stable, permitted endpoint returns the records you need, reproduce that request with Python or Scrapy instead of rendering every page. This hybrid approach gives you browser-level fidelity during discovery and lower overhead during collection.

This guide shows a complete workflow for JavaScript-heavy single-page applications (SPAs): installing Playwright, waiting for rendered content, observing network traffic, extracting data, switching to direct HTTP requests, handling pagination and authentication, and diagnosing failures. Review the target site’s robots.txt, terms of service, rate limits, and access controls before collecting anything.

1. Choose browser rendering or a direct request

Approach Use it when Main trade-off
Playwright browser JavaScript rendering, clicks, scrolling, client-side computation, login flows, or endpoint discovery is required Browser startup and memory cost; more moving parts
Selenium Your organization already has Selenium infrastructure or WebDriver-specific integrations More verbose synchronization and network inspection in many workflows
Direct HTTP request A stable, allowed API request returns the required data You must reproduce headers, cookies, parameters, pagination, and authentication correctly
Hybrid You need a browser to discover or authorize the request, then want efficient collection Requires maintaining both discovery and request code

Scrapy’s dynamic-content guidance recommends reproducing the data-bearing request when a page fetches its data separately. Keep the browser for the cases where rendering or interaction is essential.

2. Install Playwright for Python

Playwright supports Chromium, Firefox, and WebKit. Its Python installation installs the package first, then downloads browser binaries. Browsers run headlessly by default. See the official Python guide.

python -m venv .venv
source .venv/bin/activate  # Windows: .venv\\Scripts\\activate
python -m pip install --upgrade pip
pip install playwright
playwright install chromium

Install all supported browser engines only when you need cross-browser behavior:

playwright install

3. Render an SPA and extract visible content

Do not treat page.goto() completing as proof that the application is ready. The initial document may be only a shell. Wait for a meaningful selector, URL transition, or response that represents the data you need.

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/products"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(
        locale="en-US",
        timezone_id="UTC",
        viewport={"width": 1440, "height": 1000},
    )
    page = context.new_page()
    page.set_default_timeout(15_000)

    try:
        response = page.goto(URL, wait_until="domcontentloaded", timeout=30_000)
        if response is None:
            raise RuntimeError("Navigation returned no response")
        if response.status >= 400:
            raise RuntimeError(f"Navigation failed with HTTP {response.status}")

        page.locator("[data-testid='product-card']").first.wait_for(state="visible")
        products = page.locator("[data-testid='product-card']").evaluate_all(
            "cards => cards.map(card => ({" 
            "name: card.querySelector('[data-name]')?.textContent?.trim()," 
            "price: card.querySelector('[data-price]')?.textContent?.trim()" 
            "}))"
        )
        print(products)
    except PlaywrightTimeoutError:
        page.screenshot(path="timeout-debug.png", full_page=True)
        print("The expected application state was not reached")
    finally:
        context.close()
        browser.close()

Use a browser context deliberately

A context isolates cookies and storage. Set only the controls your job needs:

context = browser.new_context(
    user_agent="MyResearchBot/1.0 (+https://example.org/contact)",
    locale="en-GB",
    timezone_id="Europe/London",
    geolocation={"latitude": 51.5072, "longitude": -0.1276},
    permissions=["geolocation"],
    viewport={"width": 1280, "height": 900},
    java_script_enabled=True,
)

Do not use a custom user agent to misrepresent identity or bypass restrictions. Persist authentication only when the site permits it and the account is authorized:

context = browser.new_context(storage_state="state.json")
# After an authorized login flow:
context.storage_state(path="state.json")

4. Synchronize on application state

Fixed sleeps are a weak primary synchronization method because network and rendering times vary. Prefer one of these signals:

  • A selector appears or becomes visible.
  • The URL changes after a client-side route transition.
  • A response with a known URL or status arrives.
  • A page-specific condition becomes true.
# Selector readiness
page.locator("main [data-loaded='true']").wait_for()

# URL readiness
page.wait_for_url("**/results?query=python")

# Application condition
page.wait_for_function("""() => window.__APP_READY__ === true""")

Register response listeners before the click or navigation that triggers the request. Playwright documents request lifecycle events including request, response, requestfinished, and requestfailed.

with page.expect_response(
    lambda response: "/api/products" in response.url and response.request.method == "GET",
    timeout=20_000,
) as response_info:
    page.get_by_role("button", name="Load products").click()

api_response = response_info.value
if api_response.status >= 400:
    raise RuntimeError(f"API returned {api_response.status}")
data = api_response.json()

5. Inspect XHR and fetch traffic

Playwright can monitor and modify HTTP and HTTPS traffic, including XHR and fetch requests. Logging the method, URL, status, and resource type helps you identify the request that carries the records.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()

    def log_request(request):
        if request.resource_type in {"xhr", "fetch"}:
            print("REQUEST", request.method, request.url)

    def log_response(response):
        if response.request.resource_type in {"xhr", "fetch"}:
            print("RESPONSE", response.status, response.url)

    page.on("request", log_request)
    page.on("response", log_response)
    page.goto("https://example.com/app", wait_until="domcontentloaded")
    page.get_by_role("button", name="Search").click()
    page.wait_for_timeout(1000)
    browser.close()

For a request you have identified, capture its body while checking status and content type:

def is_data_response(response):
    return (
        response.request.resource_type in {"xhr", "fetch"}
        and "/api/" in response.url
        and response.status == 200
    )

with page.expect_response(is_data_response) as info:
    page.get_by_role("button", name="Next page").click()

response = info.value
content_type = response.headers.get("content-type", "")
if "json" not in content_type:
    raise RuntimeError(f"Unexpected content type: {content_type}")
payload = response.json()

Record enough details to reproduce the request

  • HTTP method and complete URL, including query parameters.
  • Request body for POST, PUT, or GraphQL operations.
  • Relevant headers such as Authorization, CSRF tokens, locale, and content type.
  • Cookies and whether they are session-specific.
  • Status code, response headers, and JSON shape.
  • Pagination fields such as cursor, offset, page size, and total count.

Redact tokens and personal data before storing logs. Avoid collecting secrets from pages or browser storage.

6. Reproduce the data request with Python

Once the endpoint is stable and permitted, use an HTTP client. This removes browser rendering overhead and makes retries, pagination, and parsing explicit.

import requests

API_URL = "https://example.com/api/products"

session = requests.Session()
session.headers.update({
    "Accept": "application/json",
    "User-Agent": "MyResearchBot/1.0 (+https://example.org/contact)",
})

response = session.get(
    API_URL,
    params={"query": "python", "page": 1, "page_size": 50},
    timeout=(10, 30),
)
response.raise_for_status()
payload = response.json()

for product in payload.get("items", []):
    print(product.get("name"), product.get("price"))

For a JSON POST request:

response = session.post(
    API_URL,
    json={"query": "python", "filters": {"available": True}},
    timeout=(10, 30),
)
response.raise_for_status()
data = response.json()

7. Handle pagination deterministically

Prefer a documented cursor or explicit page number. Stop when the server reports no next cursor or when the returned item count is zero. Set a hard maximum to prevent an unexpected loop.

import time
import requests

session = requests.Session()
cursor = None
seen_cursors = set()
all_items = []

for _ in range(1000):
    params = {"limit": 100}
    if cursor:
        if cursor in seen_cursors:
            raise RuntimeError("Server repeated a pagination cursor")
        seen_cursors.add(cursor)
        params["cursor"] = cursor

    response = session.get("https://example.com/api/items", params=params, timeout=30)
    response.raise_for_status()
    payload = response.json()
    items = payload.get("items", [])
    all_items.extend(items)

    cursor = payload.get("next_cursor")
    if not cursor or not items:
        break
    time.sleep(0.25)
else:
    raise RuntimeError("Pagination limit reached")

print(f"Collected {len(all_items)} items")

8. When the browser is still required

Stay with Playwright when the endpoint is not stable or accessible independently, or when the result depends on:

  • Login, consent, or a multi-step session established in the browser.
  • Client-side cryptographic or computed parameters.
  • Clicks, scrolling, virtualized lists, or lazy-loaded content.
  • Data assembled from several requests and transformed in JavaScript.
  • Rendering that changes what is available to the user.

You can still reduce work by intercepting unnecessary resources:

def route_handler(route):
    request = route.request
    if request.resource_type in {"image", "font", "media"}:
        route.abort()
    else:
        route.continue_()

page.route("**/*", route_handler)

Only block resources after confirming they are not needed for the data or application state you collect.

9. Reliability and error handling

Symptom Likely cause Fix
Empty HTML from requests Records are inserted after JavaScript runs Use Playwright or identify the data request and call it directly
Timeout waiting for a selector Wrong selector, failed API request, consent gate, or application error Save a screenshot and HTML, inspect console/network errors, and verify the selector after rendering
goto() succeeds but data is missing Navigation completion is not application readiness Wait for a meaningful selector, URL, or response
HTTP 404 or 500 appears in a response event The server returned an error response Check response.status; a completed response is not necessarily a successful one
Works manually but fails headlessly Viewport, timing, storage, permissions, or environment differs Set context options explicitly and capture diagnostics; do not bypass access controls
Direct request returns 401 or 403 Missing or expired authorization, cookies, CSRF token, or disallowed access Use an authorized session, refresh credentials, or stop if the site does not permit automation
JSON parsing fails HTML error page, redirect, rate-limit response, or changed schema Inspect status, content type, and a bounded response sample before parsing
Duplicate or missing records Unstable sorting, cursor reuse, or concurrent changes Use deterministic ordering, cursor pagination, deduplication keys, and a recorded cutoff

Capture diagnostics on failure

from pathlib import Path

try:
    page.locator("[data-testid='result']").wait_for(timeout=15_000)
except Exception:
    Path("debug.html").write_text(page.content(), encoding="utf-8")
    page.screenshot(path="debug.png", full_page=True)
    print("URL:", page.url)
    print("Title:", page.title())
    raise

Retry carefully

Retry only operations that are safe to repeat. Use a small exponential backoff for transient connection failures and selected 429 or 5xx responses. Do not retry indefinitely, and respect the site’s rate limits.

import random
import time
import requests

for attempt in range(5):
    try:
        response = requests.get("https://example.com/api/items", timeout=30)
        if response.status_code == 429 or response.status_code >= 500:
            raise requests.HTTPError(f"retryable status {response.status_code}")
        response.raise_for_status()
        payload = response.json()
        break
    except (requests.RequestException, ValueError):
        if attempt == 4:
            raise
        delay = min(30, 2 ** attempt) + random.random()
        time.sleep(delay)

10. Performance, cost, and maintenance

  • Browser startup: reuse one browser process and create isolated contexts for related jobs instead of launching a browser for every URL.
  • Parallelism: add workers gradually. More pages increase CPU, memory, bandwidth, and the chance of triggering rate limits.
  • Waiting: event-based waits usually finish sooner and fail more clearly than long fixed delays.
  • Network: direct requests are generally simpler to paginate and retry once the endpoint is known.
  • Schema drift: validate required fields and record a schema version or sample so changes are visible.
  • Data volume: stream or batch results instead of keeping an unbounded list in memory.
  • Operational cost: browser jobs consume more compute than HTTP calls; measure your own workload rather than relying on generic benchmarks.

11. Compliance and responsible collection

  • Read robots.txt and the site’s terms before collecting data.
  • Honor authentication requirements, rate limits, and explicit access restrictions.
  • Do not bypass CAPTCHAs, bot checks, paywalls, or other technical controls.
  • Collect the minimum personal data needed, protect it, and define retention and deletion rules.
  • Identify your client responsibly with a contact address where appropriate.
  • Keep an audit trail of URLs, timestamps, status codes, and consent or authorization decisions.

12. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when you need a rendered page image or PDF rather than a custom data extractor. Its capture request is a single GET:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Before capture, it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Can I scrape an SPA with only BeautifulSoup?

BeautifulSoup parses HTML already returned to Python. It cannot execute the JavaScript that populates many SPAs. Use Playwright for rendering or call the permitted data endpoint directly.

Should I choose Playwright or Selenium?

Either can drive a browser. Playwright’s Python API provides direct response waiting and network monitoring that fit SPA discovery well. Existing Selenium infrastructure, WebDriver integrations, or team experience can justify Selenium.

Is waiting for network idle enough?

Not always. Analytics, polling, and long-lived connections can prevent an idle state, while the application may still be waiting on a specific request. Prefer a selector, URL, or response tied to the data you need.

How do I know whether an API request is safe to reproduce?

Confirm that the endpoint is stable, returns the intended data, and is allowed by the site’s rules and your authorization. Preserve required authentication and rate limits, and stop when access controls prohibit the request.

Can I run this in a scheduled job?

Yes. Pin Python and Playwright versions, install browser binaries in the worker image, set explicit timeouts, persist structured logs, and alert on schema or status changes.