How to Capture XHR Responses with Playwright and SeleniumBase
Capture XHR and fetch response bodies reliably with Playwright waiters, listeners, and SeleniumBase CDP, including races and service-worker caveats.
To capture one XHR or fetch response, register the Playwright response waiter before the click or navigation that triggers it, then inspect the returned response. For continuous capture, subscribe to page.on('response') and filter by URL, method, or status. In SeleniumBase, use CDP Mode: listen for Network.ResponseReceived, keep events whose resource type is XHR, save the request ID, and call Network.getResponseBody.
The ordering prevents the most common race: the request can finish before a listener or waiter is attached. Playwright’s response event means headers and status have arrived; the body may still be downloading. SeleniumBase’s documented XHR recipe uses Chrome DevTools Protocol (CDP) and is shown as an asynchronous Python workflow.
1. Choose the capture pattern
| Need | Playwright | SeleniumBase |
|---|---|---|
| One response caused by one action | page.waitForResponse() (JavaScript) or page.expect_response() (Python) |
Register a CDP response handler, then wait for your task condition |
| Observe a stream of API traffic | page.on('response') |
Collect matching Network.ResponseReceived events |
| Read the body | Use the matched Playwright Response API after the response is available |
Call CDP Network.getResponseBody(request_id) and retain the base64 flag |
| Runtime in the cited material | JavaScript or Python, sync or async | Python async CDP example |
Playwright’s official network guides document response waiters, listeners, URL predicates, and regular-expression matching for JavaScript and Python. The SeleniumBase workflow is documented in raw_xhr_async.py and the CDP Mode guide.
2. Playwright: capture one response in Python
Use a context manager around the action. The predicate below matches both the endpoint and HTTP method, so an unrelated request cannot satisfy the waiter.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto('https://example.com/dashboard')
with page.expect_response(
lambda response: '/api/items' in response.url
and response.request.method == 'GET'
) as response_info:
page.get_by_role('button', name='Load items').click()
response = response_info.value
print('status:', response.status)
print('url:', response.url)
print('body:', response.text())
browser.close()
The waiter is created before click(). You can match a complete URL, a glob, a regular expression, or a predicate. A predicate is usually easiest when query parameters or hostnames vary. Playwright glob patterns match the entire URL, so a partial glob that omits the rest of the URL can miss the request.
Async Python
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
await page.goto('https://example.com/dashboard')
async with page.expect_response(
lambda response: '/api/items' in response.url
and response.request.method == 'GET'
) as response_info:
await page.get_by_role('button', name='Load items').click()
response = await response_info.value
print(response.status)
print(await response.text())
await browser.close()
asyncio.run(main())
3. Playwright: capture one response in JavaScript
Start the promise without awaiting it, trigger the request, then await the saved promise.
import { chromium } from 'playwright';
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto('https://example.com/dashboard');
const responsePromise = page.waitForResponse(response =>
response.url().includes('/api/items') &&
response.request().method() === 'GET'
);
await page.getByRole('button', { name: 'Load items' }).click();
const response = await responsePromise;
console.log(response.status(), response.url());
console.log(await response.text());
await browser.close();
If several requests match, make the predicate specific: include the method, path, query parameter, or an expected status. If you intentionally need the first matching response, keep the predicate broad but understand that retries or background refreshes may be selected.
4. Playwright: capture a continuous stream
Attach a listener before navigation or before the action that generates traffic.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
def on_response(response):
if '/api/' not in response.url:
return
print(response.status, response.request.method, response.url)
# Consume the body only for responses you actually need.
try:
print(response.text())
except Exception as exc:
print('body unavailable:', exc)
page.on('response', on_response)
page.goto('https://example.com/dashboard')
page.get_by_role('button', name='Load items').click()
page.wait_for_timeout(1000)
browser.close()
A response event occurs when status and headers arrive. The request lifecycle then proceeds to body completion and requestfinished. A 404 or 503 is still a valid HTTP response; requestfailed is for a client or network-level failure. If your result requires a complete body, wait for the request to finish or consume the response body through the response API. See the Request API.
JavaScript listener
page.on('response', async response => {
if (!response.url().includes('/api/')) return;
console.log(response.status(), response.url());
try {
const body = await response.text();
console.log(body);
} catch (error) {
console.error('Could not read body', error);
}
});
5. Service workers and routing gaps
Service workers can handle requests inside the browser. In cases where native routing does not observe the traffic you expect, Playwright’s network guide recommends launching the context with service workers blocked:
context = browser.new_context(service_workers='block')
Blocking changes page behavior, so use it when your test or diagnostic goal permits that change. If you need the production service-worker path, inspect BrowserContext events and identify responses handled by a service worker. The service-worker guide explains those distinctions.
6. SeleniumBase: retrieve XHR bodies through CDP
SeleniumBase’s documented raw_xhr_async.py example uses CDP Mode. The essential sequence is:
- Start an asynchronous CDP driver.
- Register a handler for
Network.ResponseReceived. - Filter for
Network.ResourceType.XHR. - Store each response URL and request ID.
- Call
Network.getResponseBody(request_id)and retain both the body and its base64 indicator.
import asyncio
from seleniumbase import SB
from seleniumbase.undetected import cdp_driver
from seleniumbase.fixtures import constants
async def main():
driver = await cdp_driver.start_async()
page = await driver.get('https://example.com/dashboard')
captured = []
async def response_handler(event):
if event.type != constants.Network.ResourceType.XHR:
return
captured.append({
'url': event.response.url,
'request_id': event.request_id,
})
await page.add_handler(constants.Network.ResponseReceived, response_handler)
# Trigger the XHR using the page's CDP/browser actions here.
await page.click('button.load-items')
# Replace this with a task-specific completion condition in production.
await asyncio.sleep(1)
for item in captured:
try:
result = await page.send(
constants.Network.get_response_body(item['request_id'])
)
item['body'] = result.body
item['base64_encoded'] = result.base64_encoded
except Exception as exc:
item['body_error'] = str(exc)
for item in captured:
print(item)
asyncio.run(main())
Match the imports and event names to the SeleniumBase version installed in your project. The sample’s quiet-period delay is a batching strategy, not a guarantee that all relevant traffic has arrived. Prefer a known UI completion signal, an application-level condition, or a bounded timeout.
7. Filter, decode, and store responses safely
Filter early
Filtering by URL and resource type reduces body reads and makes results deterministic. Add method, status, content type, or a request header when those identify the operation better than a path alone.
Preserve response metadata
Store at least the URL, request ID (for CDP), status, headers, and capture time with the body. A 404 or 503 can be useful test evidence and should not be discarded as a network failure.
Honor the base64 flag
SeleniumBase’s CDP result reports whether the body is base64 encoded. Decode it only when that flag is true; otherwise treat the value as text or bytes according to the endpoint’s content type.
Protect sensitive data
XHR bodies can contain tokens, personal data, or account information. Redact authorization headers and secrets before writing logs, and set retention limits for captured bodies.
8. Reliability and performance
- Prevent races: create waiters and listeners before the triggering action.
- Bound waits: set a timeout and report the predicate, URL, and last observed requests when it expires.
- Reduce work: filter before reading bodies; body consumption adds memory and I/O.
- Handle retries: distinguish the intended request from automatic retry or polling traffic by matching method, URL, and payload-related signals.
- Use a completion condition: a fixed sleep can miss slow traffic or waste time on fast pages.
- Keep browser versions aligned: CDP event shapes and SeleniumBase methods depend on the installed browser and package versions.
- Record failures separately: an HTTP error is a response; a client-level failure is represented by
requestfailedin Playwright.
The cited sources do not establish that Playwright or SeleniumBase is universally faster or more reliable. Measure against your target site, browser version, concurrency, and response sizes.
9. Troubleshooting
The Playwright waiter times out
Register it before the action. Check whether the predicate expects the wrong host, path, query string, method, or status. Use a temporary broad listener to print observed URLs, then narrow the predicate.
The listener sees headers but the body is empty
The response event precedes body completion. Wait for the request to finish or consume the body through the response API. Some bodies also become unavailable after the browser context closes.
A 404 or 503 appears as a response
That is expected: the server returned an HTTP response. Treat the status as application evidence; reserve network-failure handling for requestfailed.
Routing misses requests handled by a service worker
Try a context with service_workers='block' when that behavior is acceptable, or inspect BrowserContext events and service-worker response metadata.
SeleniumBase cannot retrieve a saved body
Ensure the response handler saved the request ID before calling Network.getResponseBody. Retrieve promptly, catch protocol exceptions, and keep the result’s base64 indicator. A body is not guaranteed to remain available for every timing condition.
Too many matching responses are captured
Add method and exact path checks, ignore preflight or analytics requests, and stop collecting once the task-specific completion condition is met.
The SeleniumBase example does not run as written
CDP Mode is distinct from standard WebDriver operation. Confirm the installed SeleniumBase version, use its documented CDP imports and methods, and do not assume a WebDriver method has a CDP equivalent.
10. Or skip the browser setup
If your goal is a clean visual record of a page rather than inspecting API payloads, ScreenshotNeo provides a single screenshot request. Its capture pipeline accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the result with X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for the full option set. cURL:
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
You can configure full-page capture with lazy images loaded, CSS-element capture, dark mode, device presets or custom viewports, retina scale, PDF output, custom CSS and JavaScript, clicks, waits, hidden selectors, blocked requests, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs, webhooks, bulk capture, and usage reporting. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
11. FAQ
Should I use a waiter or a listener?
Use a waiter when one action should produce one known response. Use a listener for discovery, diagnostics, or a stream of requests.
Does a 500 mean Playwright failed to capture the request?
No. A 500 is an HTTP response. Inspect its status and body; handle network-level failures separately.
Can I capture fetch as well as XHR?
Playwright response events cover browser responses regardless of whether the page initiated them with XHR or fetch. SeleniumBase’s recipe specifically filters CDP events whose resource type is XHR, so adjust the filter if you need other resource types.
Why does the SeleniumBase recipe keep a request ID?
CDP uses that ID to retrieve the corresponding body with Network.getResponseBody.
Is a fixed sleep enough after the last XHR?
It is only a sample batching technique. A page-specific completion condition and bounded timeout are more predictable.


