How to Capture Background Requests with Headless Browsers
Capture XHR and fetch traffic reliably in Playwright and Puppeteer, synchronize requests with clicks, inspect bodies, and troubleshoot missing events.
To capture background requests, attach network listeners before navigation or before the action that triggers the request. In Playwright, use page.on('request') for outgoing metadata, page.on('response') for status and headers, and page.waitForResponse() to synchronize a known API call with a click or form submission. Use routing only when you need to block, rewrite, fulfill, or abort traffic.
A reliable workflow is:
- Create the browser context.
- Attach request, response, and failure listeners.
- Navigate to the page.
- Start a waiter before clicking or submitting.
- Read a bounded response body and persist only the fields you need.
Playwright: log XHR and fetch requests
This runnable Node.js example records request metadata, response status, selected headers, and response bodies for XHR and Fetch traffic.
import { chromium } from 'playwright';
const targetUrl = 'https://example.com';
const browser = await chromium.launch();
const context = await browser.newContext({
// Use this when routing must also cover Service Worker traffic.
serviceWorkers: 'block'
});
const page = await context.newPage();
const requestIds = new WeakMap();
let nextId = 1;
page.on('request', request => {
const id = nextId++;
requestIds.set(request, id);
const type = request.resourceType();
if (type === 'xhr' || type === 'fetch') {
console.log(JSON.stringify({
event: 'request',
id,
timestamp: new Date().toISOString(),
method: request.method(),
url: request.url(),
resourceType: type,
headers: request.headers()
}));
}
});
page.on('response', async response => {
const request = response.request();
const type = request.resourceType();
if (type !== 'xhr' && type !== 'fetch') return;
const id = requestIds.get(request);
const record = {
event: 'response',
id,
timestamp: new Date().toISOString(),
status: response.status(),
url: response.url(),
headers: response.headers()
};
// Keep body capture bounded and tolerate non-JSON responses.
try {
const body = await response.text();
record.body = body.slice(0, 100_000);
} catch (error) {
record.bodyError = String(error);
}
console.log(JSON.stringify(record));
});
page.on('requestfailed', request => {
const type = request.resourceType();
if (type === 'xhr' || type === 'fetch') {
console.log(JSON.stringify({
event: 'requestfailed',
id: requestIds.get(request),
url: request.url(),
error: request.failure()?.errorText ?? 'unknown'
}));
}
});
await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
await page.waitForLoadState('networkidle').catch(() => {});
await browser.close();
Playwright’s request lifecycle for a successful request is request, response, then requestfinished. A server response such as 404 or 503 is still a response; it is not a transport failure. Use requestfailed for DNS errors, connection resets, timeouts, and similar failures.
Capture the request triggered by a click
Arm waitForResponse before the action. Starting it afterward creates a race in which a fast request has already completed.
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/data') &&
response.request().method() === 'GET' &&
response.status() === 200
);
await page.getByRole('button', { name: 'Load data' }).click();
const response = await responsePromise;
const contentType = response.headers()['content-type'] ?? '';
const body = contentType.includes('application/json')
? await response.json()
: await response.text();
console.log({ url: response.url(), status: response.status(), body });
waitForResponse accepts a glob, regular expression, or predicate. Match the URL, HTTP method, and an identifying query parameter or response status so an unrelated request cannot satisfy the waiter.
Form submission example
const responsePromise = page.waitForResponse(r =>
new URL(r.url()).pathname === '/api/search' &&
r.request().method() === 'POST'
);
await page.getByLabel('Search').fill('headless browsers');
await page.getByRole('button', { name: 'Search' }).click();
const response = await responsePromise;
console.log(await response.json());
Capture requests and responses in Python
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
context = browser.new_context(service_workers="block")
page = context.new_page()
def on_request(request):
if request.resource_type in ("xhr", "fetch"):
print("->", request.method, request.url)
def on_response(response):
request = response.request
if request.resource_type not in ("xhr", "fetch"):
return
print("<-", response.status, response.url)
try:
print(response.text()[:100000])
except Exception as exc:
print("body unavailable:", exc)
page.on("request", on_request)
page.on("response", on_response)
page.goto("https://example.com", wait_until="domcontentloaded")
browser.close()
Puppeteer equivalent
Passive response listeners work without interception. Enable interception only when you must modify traffic. Once interception is enabled, every intercepted request stalls until your handler calls continue(), respond(), or abort().
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
const page = await browser.newPage();
page.on('request', request => {
if (['xhr', 'fetch'].includes(request.resourceType())) {
console.log('->', request.method(), request.url());
}
});
page.on('response', async response => {
const request = response.request();
if (!['xhr', 'fetch'].includes(request.resourceType())) return;
console.log('<-', response.status(), response.url());
try {
console.log((await response.text()).slice(0, 100000));
} catch (error) {
console.log('body unavailable:', String(error));
}
});
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await browser.close();
Modify or block traffic with Puppeteer
await page.setRequestInterception(true);
page.on('request', request => {
if (request.isInterceptResolutionHandled()) return;
if (request.resourceType() === 'image') return request.abort();
return request.continue();
});
Playwright routing: block, rewrite, or fulfill
Use page.route() for one page and browserContext.route() when every page in a context must be covered. Define routes before navigation. Page routes take precedence over context routes when both match.
await context.route('**/analytics/**', route => route.abort());
await context.route('**/api/data', async route => {
const upstream = await route.fetch();
const json = await upstream.json();
json.debug = true;
await route.fulfill({ response: upstream, json });
});
A matching route pauses the request. Always call exactly one of route.continue(), route.fulfill(), or route.abort(), including in error paths.
Why requests go missing
- Listener registered too late: attach listeners before
goto(), clicking, typing, or submitting. - Service Worker handled the request: page and context routing do not intercept requests handled by a Service Worker. Create the context with
serviceWorkers: 'block'when complete page-level coverage matters. - Wrong resource type: GraphQL and API calls may appear as
fetch,xhr, or occasionallydocumentafter a navigation. Log all types while diagnosing. - Redirects and retries: correlate request and response objects; do not deduplicate by URL alone.
- Cross-origin restrictions: the browser may send an OPTIONS preflight before the actual request. Record both and identify the application request by method and URL.
- Response body already consumed: read the body once or store it immediately. Some failures, downloads, opaque responses, and aborted requests have no readable body.
- Long polling or WebSockets: a page may never become idle. Use a targeted waiter and an explicit timeout instead of waiting forever for network idle.
- Cache hits: a cached resource can produce different timing and event behavior than a network fetch. Disable cache only when reproducing a cache-sensitive bug.
What to record safely
For each relevant request, keep a request ID, timestamp, URL, method, resource type, selected request headers, selected response headers, status, redirect relationship, and a bounded body. Redact cookies, authorization headers, access tokens, and personal data before writing logs. Store request and response IDs together so redirects and retries are not mistaken for duplicate API calls.
function safeHeaders(headers) {
const blocked = new Set(['cookie', 'set-cookie', 'authorization', 'proxy-authorization']);
return Object.fromEntries(
Object.entries(headers).filter(([name]) => !blocked.has(name.toLowerCase()))
);
}
Efficiency, reliability, and cost
- Start with passive listeners. They do not change page behavior and are easier to debug.
- Filter by resource type, hostname, path, and method before reading bodies.
- Read only the first bounded portion of large responses, or save selected JSON fields.
- Use a specific
waitForResponsepredicate for action-triggered traffic. - Set explicit navigation and response timeouts, and record failures separately from HTTP error responses.
- Abort images, media, or stylesheets only after confirming the application does not need them for tokens, state, or rendering. An allowlist of
document,script,xhr, andfetchcan reduce work for some pages, but it is workload-dependent. - Parallelize independent pages carefully. Too much concurrency increases memory use and can trigger rate limits or change application behavior.
cURL: useful for a direct HTTP check, not browser traffic
cURL does not execute JavaScript, Service Workers, or browser interactions, so it cannot capture background requests created by a page. It is useful for checking an endpoint once you know its URL and required headers.
curl -i --compressed \
-H 'Accept: application/json' \
'https://example.com/api/data'
Or skip the browser setup
If your goal is a clean screenshot rather than inspecting every API call, ScreenshotNeo provides a single GET request. It accepts the cookie or consent banner like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets each cleanup step be turned off. Only clean shots are billed: bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with the result identified by X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options, including full-page capture, CSS element capture, device presets, dark mode, custom CSS and JavaScript, waits, blocked resources, headers, cookies, geolocation, PDFs, caching, signed links, async jobs, bulk capture, and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Troubleshooting checklist
- Confirm listeners and routes are registered before navigation or the triggering action.
- Log every resource type temporarily; then narrow the filter.
- Check whether a Service Worker owns the request and try
serviceWorkers: 'block'. - Distinguish HTTP errors from transport failures.
- Match method, URL path, and status in response waiters.
- Increase the targeted timeout for slow APIs instead of waiting indefinitely for network idle.
- Redact secrets before persisting headers or bodies.
FAQ
Can I capture a request made after a button click?
Yes. Create page.waitForResponse() first, then click the button and await the promise.
Should I use routing for logging?
No. Passive request and response listeners are sufficient for observation. Routing is for changing, blocking, fulfilling, or aborting traffic.
Why did I see a 404 in the response listener?
A 404 is a valid HTTP response. Record it as a response and reserve requestfailed for transport-level failures.
How do I capture Service Worker requests?
Page and context routes do not intercept requests handled by a Service Worker. Block Service Workers when you need page-level coverage, or use the framework’s Service Worker support when observing the worker itself.
Can cURL replace Playwright or Puppeteer?
No. cURL can call a known endpoint, but it does not run the page’s JavaScript or observe browser-generated background traffic.


