Smart Fetch Scraping: API Requests with Browser Fallbacks
Build scrapers that try fast API requests first, validate responses, and fall back to Playwright only when browser rendering is required.
Smart fetch scraping starts with the cheapest request that can work. Call the site’s API or reproduce its network request, validate the response semantically, and launch a browser only when the response is blocked, incomplete, JavaScript-only, or requires interaction. This reduces latency and browser resource use while preserving a reliable fallback for browser-dependent pages.
What smart fetch scraping means
A smart fetcher is a two-stage pipeline:
- Direct tier: send an HTTP request to the underlying API or document endpoint.
- Browser tier: if validation fails, use Playwright or a managed browser to render the page and complete browser-only steps.
Do not treat HTTP 200 as proof that extraction succeeded. A successful status can contain a login page, bot challenge, empty JavaScript shell, stale cache, or partial data. Validate the content you actually need before returning it.
Browserless describes the same cascading strategy: try a fast HTTP fetch first and launch a full browser only if the initial fetch fails or returns incomplete content. Scrapy’s dynamic-content guidance similarly recommends finding the data source and reproducing the request before resorting to a headless browser.
When to use an API request or a browser
| Question | Prefer a direct request when… | Use browser fallback when… |
|---|---|---|
| Is the data in HTML or JSON? | The response contains the complete fields or markup. | The response is only a JavaScript shell or omits the target data. |
| Does the page need JavaScript? | The API can be called independently. | Client code computes values or triggers required requests. |
| Is session state required? | You can send valid cookies, tokens, or headers. | State is created by browser navigation or an interaction. |
| Are there browser-only actions? | No clicks, scrolling, consent handling, or DOM events are needed. | The flow requires interaction, challenge handling, or browser cookies. |
| What matters most? | Lower latency, transfer, and resource use. | Rendered-page fidelity and browser behavior. |
Build the pipeline step by step
1. Define the success contract
Write down what a valid result contains before writing the fallback. Examples include a JSON field such as items, at least one article record, an HTML marker such as <main>, or a minimum number of product cards. A contract prevents a login page or empty shell from being mistaken for success.
2. Send the cheapest direct request
Use the method, URL, query, body, authentication headers, cookies, and user agent observed in the application. For a dynamic site, inspect browser network activity and reproduce the request. Scrapy documents exporting a browser request as cURL and translating it into a Scrapy request.
3. Validate status, type, schema, and completeness
- Check the status code and redirect destination.
- Check the
Content-Typeheader. - Reject known login, challenge, and error-page markers.
- Parse JSON and verify required fields and types.
- Check that the record count or HTML markers meet your contract.
- Record the validation failure as the reason for escalation.
4. Reproduce the underlying request when possible
Open browser developer tools, inspect the Network panel, identify the XHR or fetch request containing the data, and copy it as cURL. Compare its URL, method, query, request body, authorization, cookies, and required headers with your direct client. This usually returns structured data with less parsing and transfer than rendering the entire page.
5. Escalate to Playwright only when required
Use a browser for JavaScript execution, DOM events, browser-created cookies, challenge flows that your access is permitted to handle, or pages whose data cannot be reproduced as an HTTP request. Keep the browser context isolated per account or job when session leakage would be harmful.
6. Return normalized data and telemetry
Return the extracted result with fields such as tier (direct or browser), escalation_reason, elapsed time, retry count, final URL, and failure category. This lets you find endpoints that need a better direct implementation instead of silently paying browser costs forever.
Complete Python implementation
This example tries JSON first, validates the payload, then falls back to Playwright. Replace the URL and selectors with the target site’s documented or observed requests.
import json
import re
import time
from typing import Any
import requests
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
DIRECT_URL = "https://example.com/api/products"
PAGE_URL = "https://example.com/products"
def valid_payload(response: requests.Response) -> tuple[bool, str, Any | None]:
if response.status_code != 200:
return False, f"http_status_{response.status_code}", None
content_type = response.headers.get("content-type", "").lower()
if "json" not in content_type:
return False, "unexpected_content_type", None
try:
payload = response.json()
except ValueError:
return False, "invalid_json", None
if not isinstance(payload, dict) or not isinstance(payload.get("items"), list):
return False, "missing_items", None
if not payload["items"]:
return False, "empty_items", None
return True, "ok", payload
def looks_like_challenge_or_login(html: str) -> bool:
text = html.lower()
markers = ("captcha", "access denied", "verify you are human", "sign in")
return any(marker in text for marker in markers)
def smart_fetch() -> dict[str, Any]:
started = time.monotonic()
direct_reason = "not_attempted"
try:
response = requests.get(
DIRECT_URL,
headers={"Accept": "application/json", "User-Agent": "smart-fetch/1.0"},
timeout=20,
)
ok, direct_reason, payload = valid_payload(response)
if ok:
return {
"tier": "direct",
"data": payload,
"escalation_reason": None,
"elapsed_ms": round((time.monotonic() - started) * 1000),
}
except requests.RequestException as exc:
direct_reason = f"request_error:{type(exc).__name__}"
with sync_playwright() as pw:
browser = pw.chromium.launch(headless=True)
context = browser.new_context()
page = context.new_page()
try:
page.goto(PAGE_URL, wait_until="domcontentloaded", timeout=45_000)
page.wait_for_selector("main", timeout=15_000)
html = page.content()
if looks_like_challenge_or_login(html):
raise RuntimeError("challenge_or_login_page")
cards = page.locator("[data-product]").all_inner_texts()
if not cards:
raise RuntimeError("empty_browser_result")
return {
"tier": "browser",
"data": {"products": cards},
"escalation_reason": direct_reason,
"elapsed_ms": round((time.monotonic() - started) * 1000),
}
except (PlaywrightTimeoutError, RuntimeError) as exc:
return {
"tier": "failed",
"data": None,
"escalation_reason": direct_reason,
"failure": str(exc),
"elapsed_ms": round((time.monotonic() - started) * 1000),
}
finally:
browser.close()
if __name__ == "__main__":
print(json.dumps(smart_fetch(), indent=2))
Complete Node.js implementation
import { chromium } from "playwright";
const apiUrl = "https://example.com/api/products";
const pageUrl = "https://example.com/products";
function validatePayload(response, payload) {
const type = response.headers.get("content-type") || "";
if (!response.ok) return [false, `http_status_${response.status}`];
if (!type.includes("json")) return [false, "unexpected_content_type"];
if (!payload || !Array.isArray(payload.items) || payload.items.length === 0) {
return [false, "missing_or_empty_items"];
}
return [true, "ok"];
}
async function smartFetch() {
const started = Date.now();
let reason = "not_attempted";
try {
const response = await fetch(apiUrl, {
headers: { accept: "application/json", "user-agent": "smart-fetch/1.0" },
signal: AbortSignal.timeout(20_000),
});
let payload = null;
try { payload = await response.json(); } catch { reason = "invalid_json"; }
const [valid, validationReason] = validatePayload(response, payload);
if (valid) {
return { tier: "direct", data: payload, escalation_reason: null,
elapsed_ms: Date.now() - started };
}
reason = validationReason;
} catch (error) {
reason = `request_error:${error.name}`;
}
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
try {
await page.goto(pageUrl, { waitUntil: "domcontentloaded", timeout: 45_000 });
await page.locator("main").waitFor({ timeout: 15_000 });
const html = await page.content();
if (/captcha|access denied|verify you are human|sign in/i.test(html)) {
throw new Error("challenge_or_login_page");
}
const products = await page.locator("[data-product]").allTextContents();
if (!products.length) throw new Error("empty_browser_result");
return { tier: "browser", data: { products }, escalation_reason: reason,
elapsed_ms: Date.now() - started };
} catch (error) {
return { tier: "failed", data: null, escalation_reason: reason,
failure: error.message, elapsed_ms: Date.now() - started };
} finally {
await browser.close();
}
}
console.log(await smartFetch());
Direct HTTP examples
cURL
curl -i -H 'Accept: application/json' \
-H 'User-Agent: smart-fetch/1.0' \
'https://example.com/api/products'
Python
import requests
r = requests.get(
"https://example.com/api/products",
headers={"Accept": "application/json"},
timeout=20,
)
r.raise_for_status()
payload = r.json()
assert payload.get("items"), "incomplete payload"
print(payload)
Node.js
const res = await fetch('https://example.com/api/products', {
headers: { accept: 'application/json' },
signal: AbortSignal.timeout(20000),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const payload = await res.json();
if (!Array.isArray(payload.items) || payload.items.length === 0) {
throw new Error('incomplete payload');
}
console.log(payload);
Sharing cookies and session state with Playwright
Playwright can issue HTTP methods through APIRequestContext. A request context obtained from a browser context shares that context’s cookie jar, so API calls and page navigation can use the same session. Use this when login state or a consent cookie is established in the browser.
import { chromium } from "playwright";
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
await page.goto("https://example.com/login");
// Complete the permitted login flow here.
const api = context.request;
const response = await api.get("https://example.com/api/account");
console.log(await response.json());
await browser.close();
Keep contexts isolated by user or job. Never log session cookies or authorization headers. If a direct request requires a token, prefer the site’s supported authentication mechanism and comply with its terms and access controls.
Request interception and observation
Playwright routing can intercept requests at page or browser-context scope. Use it to observe the API call a page makes, modify a permitted request, or fulfill a response in tests.
await page.route("**/api/**", async route => {
const request = route.request();
console.log(request.method(), request.url(), request.headers());
await route.continue();
});
await page.goto("https://example.com/products", { waitUntil: "networkidle" });
For production extraction, use interception to discover and verify requests, then reproduce stable API calls directly when practical. Avoid depending on private endpoints if the site does not authorize that access.
Fallback design for reliability
- Bound retries: retry transient network errors with exponential backoff and a maximum attempt count.
- Separate failure classes: distinguish timeout, non-2xx status, challenge page, schema mismatch, selector change, and resource exhaustion.
- Preserve evidence: record the final URL, status, content type, validation reason, and a redacted response sample.
- Set budgets: cap navigation time, total job time, concurrent browsers, and page size.
- Use stable waits: prefer a meaningful selector or network condition over arbitrary long sleeps.
- Make jobs idempotent: store a request key so a retry does not duplicate downstream writes.
- Monitor escalation rate: a sudden increase often means an API schema, authentication flow, or page layout changed.
Performance, cost, and scaling
Direct API reproduction generally uses less time, bandwidth, and memory than launching a browser. Browser fallback consumes more CPU and memory and has more failure points, so reserve it for requests that genuinely need rendering or interaction. The research sources provide qualitative guidance rather than a universal speed or success-rate benchmark; measure your own targets.
| Optimization | How it helps |
|---|---|
| Validate before parsing deeply | Stops bad responses quickly and avoids unnecessary work. |
| Cache validated direct responses | Reduces repeated API calls when freshness permits. |
| Reuse a browser only within a controlled context | Amortizes startup cost while preventing session cross-contamination. |
| Block irrelevant resources in browser jobs | Reduces transfer and rendering work when images or analytics are unnecessary. |
| Limit concurrency | Prevents exhausted browser workers and protects the target service. |
| Track tier and reason | Shows where browser capacity and engineering effort are being spent. |
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP 200 but no records | JavaScript shell, login page, or empty filtered response. | Validate schema and markers; inspect network calls; escalate if needed. |
| JSON parsing fails | HTML challenge or error page returned with a misleading status. | Check content type and redirect URL before parsing. |
| Direct request is unauthorized | Missing cookies, authorization, CSRF token, or required headers. | Reproduce the complete permitted request or establish state in a browser context. |
| Playwright navigation timeout | Slow page, blocked resource, or a page that never reaches the chosen load condition. | Use a realistic timeout, wait for a target selector, and classify the failure before retrying. |
| Selector timeout | Layout changed, content is inside a frame, or the request returned a challenge. | Inspect the DOM and frames, verify the page verdict, and update a stable selector. |
| Browser workers run out | Too much concurrency or contexts are not closed. | Close pages and contexts, cap concurrency, and enforce job timeouts. |
| Intermittent empty results | Race condition before data rendering or an unstable API response. | Wait for a meaningful selector or response, validate count, and use bounded retries. |
Or skip the browser setup
If your goal is a clean screenshot or PDF rather than structured records, ScreenshotNeo provides one GET request for a PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie and consent banners as a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Each step can be turned off.
Only clean shots are billed. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result through X-Page-Verdict and X-Billed headers.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the request options. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with the included monthly screenshots.
FAQ
Should every scraper use a browser first?
No. Start with the site’s API or the request that supplies the data. Use a browser when validation shows that rendering or interaction is required.
Is a 200 response enough to skip fallback?
No. Check content type, schema, required fields, and completeness. Login and challenge pages commonly return successful HTTP statuses.
How do I keep cookies between API and browser calls?
Use a Playwright browser context and its associated request context. They share the context cookie jar.
When should I use a managed browser service?
Use one when you need browser execution without operating Chromium workers yourself. Keep the same validation, bounded retries, telemetry, and access-control checks.
How can I reduce browser fallback frequency?
Inspect network calls, reproduce the stable data request, send all required authentication state, and maintain schema validation tests for the direct response.
Production checklist
- Define the required fields or DOM markers.
- Send the direct request with the correct method, headers, body, and permitted session state.
- Validate status, content type, schema, and completeness.
- Record why each escalation happened.
- Use bounded browser retries and explicit timeouts.
- Close Playwright pages, contexts, and browsers on every path.
- Redact cookies, tokens, and personal data from logs.
- Respect the target site’s terms, robots directives where applicable, rate limits, and access controls.
- Monitor direct success rate, browser escalation rate, latency, and failure categories.


