How to Scrape Network Requests in Browser Automation
Capture browser network requests, match responses, inspect failures, and choose between Playwright, Selenium WebDriver BiDi, and Chrome DevTools Protocol.
To capture network requests in browser automation, attach request and response listeners before navigation or the action that triggers traffic. In Playwright, use page.on('request') and page.on('response') to observe traffic, waitForResponse() to wait for one anticipated API response, and page.route() or browserContext.route() when you need to modify, fulfill, or abort requests. In Selenium, use WebDriver BiDi network events when your browser and language binding support the needed operations. Use Chrome DevTools Protocol (CDP) when you need its lower-level Chrome Network domain.
Only inspect traffic you are authorized to access. A response event means status and headers have arrived; it does not mean the response body has finished downloading. An HTTP 404 or 503 is still a response. A request failure means the browser did not receive an HTTP response, for example because of a network error or timeout.
1. Choose what to capture
“Scrape network requests” can mean several different things. Decide what data and control you need before choosing an API.
| Need | What to collect or do | Typical choice |
|---|---|---|
| List requests | Record URLs, methods, and resource types as the page loads or actions run. | Playwright request events; Selenium BiDi events where supported. |
| Associate HTTP results | Match each response to its URL and inspect status and headers. | Response events, or wait for a matching response. |
| Read a response body | Wait until the request completes, then use the framework’s body-reading API. | Playwright response body methods; CDP Network body methods. |
| Wait for one API call | Start a matching response wait before the action, then await it. | Playwright waitForResponse(). |
| Change or stop traffic | Modify a request, supply a response, or abort it. | Playwright routing, Selenium BiDi or CDP operations where supported. |
For passive logging, event listeners are usually the simplest approach. For one known endpoint, use a response wait with a predicate that identifies the request. For interception, choose a routing or protocol API deliberately: intercepted requests may pause until your handler resolves them.
2. Playwright: log requests and responses
Attach listeners before goto() if you want to capture navigation traffic. This runnable Node.js example logs each request’s method and URL, and each response’s status and URL.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
const page = await browser.newPage();
page.on('request', request => {
console.log('>>', request.method(), request.url());
});
page.on('response', response => {
console.log('<<', response.status(), response.url());
});
await page.goto('https://example.com');
await browser.close();
})();
Install Playwright using the instructions for your project and browser. These page-level events are useful for observing traffic from that page. If a new tab or popup makes the request, attach listeners to that page too, or listen for new pages on the browser context and then attach handlers.
Log method, URL, resource type, and response status
page.on('request', request => {
console.log({
method: request.method(),
url: request.url(),
resourceType: request.resourceType()
});
});
page.on('response', response => {
console.log({
status: response.status(),
url: response.url()
});
});
Keep request and response handling separate. A request event is useful for identifying outgoing traffic; a response event provides the HTTP result once status and headers arrive. To read a response body, wait for the request to finish and use the response body API supported by your Playwright version. For example, after a matched response resolves, call await response.text() for text or await response.json() for JSON. Handle parse errors and non-JSON responses explicitly.
3. Playwright: wait for one request or response
When an action triggers a known API call, create the wait promise before clicking. Otherwise, a fast request can start before the listener is registered.
const responsePromise = page.waitForResponse('**/api/fetch_data');
await page.getByText('Update').click();
const response = await responsePromise;
console.log('status:', response.status());
console.log('url:', response.url());
if (response.ok()) {
const body = await response.json();
console.log(body);
}
Playwright’s glob pattern must match the entire URL. The pattern **/api/fetch_data can match across path separators, but consider the protocol, host, query string, and path when the page can call similar endpoints. For complex conditions, use a predicate:
const responsePromise = page.waitForResponse(response => {
const request = response.request();
const url = new URL(response.url());
return url.pathname === '/api/fetch_data' &&
request.method() === 'POST';
});
await page.getByText('Update').click();
const response = await responsePromise;
console.log(response.status(), response.url());
Use page.waitForRequest() when you only need to confirm that the browser issued a request, rather than waiting for an HTTP response. Add a timeout appropriate to the application’s behavior so an absent request does not hang the automation indefinitely. Check the Playwright Events documentation for the current event and wait APIs.
4. Understand request, response, and failure events
The lifecycle helps distinguish server errors from transport failures:
- Request issued: the browser starts a network request.
- Response received: status and headers arrive. The body may still be downloading.
- Request finished: the response body download completes.
- Request failed: the client did not get an HTTP response, such as from a network error or timeout.
In Playwright, a 404 or 503 normally appears as a response with that status; it is not a requestfailed event. Treat HTTP status handling separately from transport failures. For APIs where an error status is a valid outcome, inspect response.status() rather than assuming every received response is successful.
For body inspection, do not try to read it at the request event. Use the matching response and the framework’s response body methods after the response is available. The body can be large, binary, or invalid JSON, so avoid logging sensitive or unnecessary content.
5. Intercept requests with Playwright routing
Use event listeners for observation. Use routing when you need to alter behavior. The route handler must resolve a matching request by continuing it, fulfilling it, or aborting it; otherwise, the request can remain stalled.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
await context.route('**/api/feature-flags', async route => {
await route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({ enabled: true })
});
});
await page.goto('https://example.com');
await browser.close();
})();
Route at the page level with page.route() when the rule should apply to one page, or at the context level with browserContext.route() when it should cover pages in that context. Check the current Playwright Network documentation for route options and request handling methods. Keep routes narrow: broad patterns can accidentally affect scripts, images, or unrelated API calls.
Playwright documents that requests handled by routing stall until the handler takes action. Always continue, fulfill, or abort on every path through the handler, including error paths. Also account for Service Workers: native page and context routing may appear to miss traffic controlled by a Service Worker. If appropriate for the test, create the context with serviceWorkers: 'block'. If the goal is to observe Service Worker requests themselves, use Playwright’s dedicated Service Worker guidance instead.
6. Selenium WebDriver BiDi, CDP, or Playwright?
| Option | Best fit | Limits to check |
|---|---|---|
| Playwright events and routing | Playwright projects needing page/context request and response events, a single-response wait, or request interception. | Service Worker behavior, popup pages, and the exact routing features supported by the installed version. |
| Selenium WebDriver BiDi | Selenium projects that need bidirectional event streaming, including network events, in a supported browser and binding. | BiDi support is evolving; check the browser, Selenium version, language binding, and operation coverage. |
| Chrome DevTools Protocol | Chrome-specific automation that needs direct access to the Chrome Network domain, including protocol-level operations. | CDP is Chrome-specific and lower-level; it is not a cross-browser standard. |
Choose based on the framework already in use, the deployed browser versions, whether you need passive events or interception, and whether you need response bodies. Selenium describes WebDriver BiDi as the bidirectional direction for event streaming, while its CDP support is temporary as BiDi implementation progresses. That does not guarantee that every BiDi operation is available in every browser or binding today.
Selenium WebDriver BiDi setup
Enable the webSocketUrl capability using the mechanism exposed by your Selenium language binding, then subscribe to the network events that the target browser and binding support. The exact event types and helper APIs are version-dependent, so use the official Selenium WebDriver BiDi documentation for code that matches your binding and release. Confirm that the session actually exposes a BiDi WebSocket before relying on network events.
BiDi is a W3C protocol intended for bidirectional browser communication. Its network event coverage and commands are still being implemented across browser vendors. If a required operation is missing, check browser support and binding support separately; a protocol specification alone does not mean a deployed browser implements every part.
Chrome DevTools Protocol Network domain
CDP exposes Chrome’s Network domain with request and response events and methods for tasks such as retrieving request post data or response bodies. Enable the Network domain and listen for the relevant events using the CDP client integrated with your automation stack. CDP gives direct protocol access, but ties the implementation to Chrome and protocol versions. Consult the official Chrome DevTools Protocol Network reference for method names, event payloads, and compatibility details.
7. Troubleshooting missing or unexpected traffic
| Symptom | Likely cause | Fix |
|---|---|---|
| The navigation request is missing from the log. | Listeners were attached after navigation started. | Register listeners before goto() or the action that causes the request. |
| A click-triggered request is missed intermittently. | The request began before the wait or listener was registered. | Create the waitForRequest() or waitForResponse() promise before clicking. |
| The response wait times out. | The endpoint pattern does not match the full URL, the action did not issue the request, or the request came from another page. | Log request URLs first, tighten or correct the predicate, verify the action, and check popup pages. |
| A 404 or 503 is reported as a failure. | HTTP errors and request failures are being conflated. | Inspect response status for HTTP errors. Reserve request-failure handling for cases where no HTTP response arrived. |
| The response status appears, but the body is unavailable. | The body is not ready at the response event, or the body was consumed/handled incorrectly. | Wait for completion as required by the framework and use its response body method. Check whether the response is binary or empty. |
| Routing does not see requests handled by a Service Worker. | The Service Worker owns the request path. | For tests where blocking is acceptable, set serviceWorkers: 'block'. Otherwise use dedicated Service Worker handling. |
| Some traffic appears in a new tab but not the original page. | Listeners are only attached to the original page. | Subscribe to context page creation and attach listeners to each new page. |
| Selenium BiDi has no network events. | The browser, driver, Selenium version, or binding lacks the needed BiDi support, or the session did not enable its WebSocket. | Enable webSocketUrl, verify the session capability and check support for the exact event and operation. |
| CDP calls fail after a browser update. | The protocol client and Chrome version may not align. | Check the matching CDP protocol documentation and update the client integration as needed. |
| Requests remain pending after interception. | A route handler did not resolve every matching request. | Ensure all handler branches call continue, fulfill, or abort, including exceptions. |
8. Performance, reliability, and data handling
- Keep collection selective. A listener for every request can generate a large log on pages with many assets or polling calls. Filter by hostname, path, method, or resource type when the framework allows it.
- Do not block the event loop with logging. Avoid synchronous disk writes and large body dumps inside event handlers. Queue only the fields needed and persist them outside the critical path.
- Set explicit waits and cleanup. Give response waits finite timeouts, close pages and browser sessions in cleanup paths, and remove listeners when a long-lived page no longer needs them.
- Separate status from transport outcome. Record HTTP status for received responses and a distinct failure reason for requests that got no response.
- Protect captured data. URLs, headers, cookies, post data, and response bodies may contain credentials, personal data, or tokens. Redact secrets, minimize retention, and use authorized test data.
- Prefer the least complex supported interface. High-level framework events are easier to maintain for ordinary logging; use routing or CDP only when you need their control or protocol-specific features.
There is no universal performance cost for network logging: it depends on request volume, body sizes, serialization, and what the handler does. Recording only metadata is generally less work than buffering and serializing every body. Measure in the target environment if capture overhead matters to test timing.
9. Or skip the browser setup
If the job is to capture a clean screenshot of a page rather than inspect its underlying network traffic, ScreenshotNeo is a website screenshot API and MCP server. It does not expose a browser network log; it provides a screenshot or PDF from one GET request. Its cookie and consent handling accepts banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and each response identifies the page verdict and billing status in headers. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.
Install the Python dependency with python -m pip install requests, then run:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as output:
output.write(r.content)
See the ScreenshotNeo API documentation for options and setup. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.
10. Frequently asked questions
Can I capture requests that happened before my listener was registered?
Typically, event listeners report events from the point they are attached. Register them before navigation or the triggering action. For a request already in progress, use the framework’s available request or response state APIs if provided rather than expecting a past event to replay.
Should I log request bodies and headers?
Only when your authorized debugging task requires them. Headers and bodies can contain secrets and personal data; redact values and minimize storage.
Is CDP the same thing as Selenium BiDi?
No. CDP is Chrome’s DevTools protocol. WebDriver BiDi is a W3C protocol with support evolving across browsers and bindings.
Can ScreenshotNeo replace network request scraping?
No. ScreenshotNeo captures page images or PDFs and provides page information through its tools; it is not a general-purpose network traffic inspector.


