ScreenshotNeo

BlogHow-to

How to Handle Events and Promises in Web Scraping

Learn how to coordinate browser events, actions, and readiness checks in Playwright and Puppeteer, with runnable JavaScript examples and fixes for common races.

By the ScreenshotNeo team4 October 202612 min read

In web scraping, handle an event-triggering action by creating the event-wait promise first, performing the action second, and awaiting the promise third. This prevents fast events from firing before your scraper is listening. Then wait for the specific outcome your extraction needs—a matching response, URL, popup, or ready element—instead of assuming that a click or general network silence means the page is ready.

This guide uses JavaScript with Playwright and Puppeteer. The examples assume their current documented APIs; check the documentation for the version installed in your project: Playwright Page API, Playwright pages and popups, and Puppeteer Page API.

1. The core pattern: wait before triggering

A browser action and the event it causes are two separate asynchronous operations. Start the wait without awaiting it, perform the action, then await the saved promise:

const eventPromise = page.waitForEvent('download');
await page.getByRole('button', { name: 'Download' }).click();
const download = await eventPromise;

If you await the event before triggering it, the script hangs waiting for something that cannot happen yet. If you click first and register afterward, a quick response or popup can be missed. Playwright documents this wait-before-action ordering for events such as downloads, requests, and popups. See its Page API and pages guide.

2. Set up a runnable Playwright scraper

Install Playwright and its Chromium browser, then save this as scrape.mjs. It visits a sample page, waits for the product API response before clicking, validates the response, and extracts rendered product titles. Replace the URL, button name, endpoint fragment, and result selector with those used by the site you are authorized to scrape.

npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';

const startUrl = 'https://example.com/products';
const browser = await chromium.launch({ headless: true });

try {
  const page = await browser.newPage();
  page.setDefaultTimeout(10_000);
  page.setDefaultNavigationTimeout(20_000);

  await page.goto(startUrl, { waitUntil: 'domcontentloaded' });

  // Subscribe before clicking. Narrow the predicate to avoid matching unrelated traffic.
  const responsePromise = page.waitForResponse(response => {
    const request = response.request();
    return response.url().includes('/api/products')
      && request.method() === 'GET';
  }, { timeout: 15_000 });

  await page.getByRole('button', { name: 'Load products' }).click();
  const response = await responsePromise;

  if (!response.ok()) {
    throw new Error(`Product API returned HTTP ${response.status()}`);
  }

  const payload = await response.json();

  // Wait for the rendered state your extraction needs, not just the API response.
  const titles = await page.locator('.product-card h2').allTextContents();
  console.log(JSON.stringify({ payload, titles }, null, 2));
} catch (error) {
  console.error('Scrape failed:', error.message);
  process.exitCode = 1;
} finally {
  await browser.close();
}

The API response and rendered DOM are distinct readiness conditions. If your data comes from the JSON response, parse that response. If you need page-rendered text, wait for the relevant locator and extract from it. In a real scraper, an endpoint might require a POST request or have a different method, query string, or response shape; adapt the predicate and extraction accordingly.

3. Choose the event or readiness condition that matches the task

What the scraper needs Use Notes
A particular API response page.waitForResponse(predicate) Match a distinctive URL and, when useful, method or status. Parse the response body after it resolves.
A request being sent page.waitForRequest(predicate) Useful to inspect outgoing method, URL, or payload; it does not establish that the server succeeded.
A known destination URL Playwright page.waitForURL() Use a glob, regular expression, URL pattern, or predicate for the expected destination.
A click-triggered navigation in Puppeteer Promise.all([page.waitForNavigation(), action]) Starts both operations together so the navigation wait is already active when the click runs.
A popup caused by a known action Playwright page.waitForEvent('popup') Register before the click, then wait on the returned popup page for the content needed.
A new page caused by an unknown action Playwright browserContext.on('page', ...) Context scope observes pages across the context; remove the listener when done.
An element becoming available or actionable Playwright locator action or web-first assertion Locators wait for action preconditions. Assert the state that matters to extraction.
Document lifecycle milestone DOMContentLoaded or load Use only when that milestone fits the task. Client-rendered content may arrive later.

Playwright describes waitForNavigation as inherently racy and recommends waitForURL for URL-based waits. Puppeteer documents pairing its navigation wait and action using Promise.all. Use the approach recommended for your library rather than assuming these similarly named APIs behave identically. Sources: Playwright Page API and Puppeteer Page API.

4. Handle common scraping events

Wait for a specific API response in Playwright

const responsePromise = page.waitForResponse(response => {
  const request = response.request();
  return new URL(response.url()).pathname === '/api/products'
    && request.method() === 'GET'
    && response.status() === 200;
}, { timeout: 15_000 });

await page.getByRole('button', { name: 'Load products' }).click();
const response = await responsePromise;
const products = await response.json();

A matcher that is too broad may resolve on an unrelated request. Match the endpoint path, and add a method, status, query value, or other distinguishing condition when the page issues similar requests. If the site makes multiple legitimate matching calls, the first response may not be the one you need; refine the predicate or inspect the returned response.

Wait for a URL after an action in Playwright

await page.getByRole('link', { name: 'Next page' }).click();
await page.waitForURL('**/products?page=2', { timeout: 15_000 });
await page.locator('.product-card').first().waitFor({ state: 'visible' });

The locator wait checks the content condition separately from the URL transition. If a client-side application updates the URL with the History API, a URL wait is often the relevant signal. If the action does not navigate, wait for the changed element or response instead.

Wait for navigation in Puppeteer

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(20_000);
  await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });

  const [navigationResponse] = await Promise.all([
    page.waitForNavigation({ waitUntil: 'domcontentloaded', timeout: 15_000 }),
    page.locator('a.next-page').click(),
  ]);

  // A History API navigation can yield null; verify the outcome you need.
  console.log('Current URL:', page.url());
  console.log('Navigation response:', navigationResponse?.status() ?? 'no document response');
  await page.locator('.product-card').wait();
  console.log(await page.locator('.product-card h2').allTextContents());
} finally {
  await browser.close();
}

Puppeteer’s locator API is shown here; if your installed version uses a different interaction API, use its documented equivalent. A navigation wait can resolve with no main-resource response for History API URL changes or anchor navigation. Verify the URL or page state as appropriate. See the Puppeteer navigation API.

Capture a popup opened by a known action

const popupPromise = page.waitForEvent('popup', { timeout: 10_000 });
await page.getByRole('link', { name: 'Open details' }).click();
const popup = await popupPromise;

await popup.waitForLoadState('domcontentloaded');
await popup.locator('main h1').waitFor({ state: 'visible' });
console.log(await popup.title());

For new pages that may be opened by an action you cannot predict, observe the browser context’s page event instead. Use page-level popup waits for a known opener and context-level page events for broader observation. See Playwright’s popup guidance.

Use persistent listeners for an event stream

A wait method is suited to a known next event and returns that event’s data. A persistent listener is suited to events that can occur repeatedly or at varying times. Remove it when it is no longer needed:

const onResponse = response => {
  if (response.url().includes('/api/products')) {
    console.log('Product response:', response.status(), response.url());
  }
};

page.on('response', onResponse);
try {
  await page.goto('https://example.com/products');
  // Perform the remaining scrape work.
} finally {
  page.off('response', onResponse);
}

Attach a one-off listener when you need only the next occurrence and the library offers one; otherwise, use a wait method with a predicate and timeout. Avoid leaving broad listeners attached across pages or jobs: they can process unrelated future events and retain references longer than necessary. Playwright documents both waiting and listener-based event handling in its events guide.

5. Pick a useful readiness signal

  1. Start with the extraction dependency. If the required data is in an API response, wait for that response. If it is rendered in the page, wait for a selector or assert the text or state needed.
  2. Use URL waits for URL outcomes. In Playwright, use waitForURL rather than deprecated, racy waitForNavigation when checking a destination.
  3. Use lifecycle events only for lifecycle needs. DOMContentLoaded means the document was parsed; load waits for the load event. Neither alone proves that a client-rendered list or delayed API result is ready.
  4. Avoid fixed sleeps as the main synchronization method. A delay can be too short on a slow run and waste time on a fast one. Prefer an observable condition. This is a practical engineering recommendation; see the library guidance on locator waits and readiness assertions.
  5. Do not use network idle as a universal signal. Playwright marks networkidle discouraged for readiness in tests and recommends web assertions instead. Pages with long-lived connections, polling, or late updates can make network silence a poor proxy for useful content. The latter cases are examples of why a condition-specific wait is more robust.

Playwright locator actions wait for their preconditions, and its documentation recommends web-first assertions to assess readiness. See the Page API and Puppeteer page interactions guide.

6. Timeouts, errors, and cancellation

Give each wait a timeout that reflects the operation, and make the timeout error identify the condition. Playwright’s event waits can fail if the page closes before the event occurs; wait options support timeouts, and current documentation also describes abort signals. Puppeteer event waits likewise support timeouts, with defaults configurable on the page. Avoid disabling timeouts unless you have another explicit cancellation mechanism.

try {
  const responsePromise = page.waitForResponse(
    response => response.url().includes('/api/products'),
    { timeout: 12_000 },
  );

  await page.getByRole('button', { name: 'Load products' }).click();
  const response = await responsePromise;
  if (!response.ok()) throw new Error(`HTTP ${response.status()}`);
} catch (error) {
  console.error('Waiting for the products response failed:', error.message);
  throw error;
}

When an action and its wait run concurrently, a failure in either can make the operation fail. Catch errors near the step that owns the wait, add the URL and expected condition to logs, and close browser resources in a finally block. For long-running workers, also make sure the job runner can stop or retry a timed-out task.

7. Troubleshooting event and promise failures

Symptom Likely cause Fix
Wait times out although the click worked The expected event did not occur, the matcher is wrong, or the action had a different outcome. Log the current URL and relevant request/response events. Confirm the button behavior, URL, endpoint path, and method. Wait for the actual resulting condition.
The event sometimes gets missed The listener or wait was registered after the trigger. Create the wait promise first, then click or navigate, then await the promise.
The wrong response resolves the wait The predicate matches analytics, prefetch, or another background request. Constrain by pathname, method, query parameters, status, or a request property that distinguishes the target call.
Navigation wait hangs after a click The click triggered an in-page update or request, not a document navigation. Wait for the expected response, URL update, or element state. Use Playwright waitForURL for URL changes.
Navigation completes but extracted content is empty The document lifecycle completed before client rendering or the relevant response. Wait for the target locator, a meaningful assertion, or the data response before extracting.
Popup wait times out The action was blocked, opened a page outside the expected scope, or did not open a popup. Verify the action and popup behavior. For a known opener, wait on that page before clicking; for unknown new pages, observe the context.
Wait fails because the page closed The tab or browser context closed before the event arrived. Check page lifecycle and cleanup ordering. Do not close the page until dependent waits and extraction complete.
Wait is slow or flaky with networkidle Network silence is not the same as extraction readiness, and some pages keep network activity open. Replace the general idle wait with a specific response, URL, or rendered-state condition.
Repeated events appear in later jobs A persistent listener was left attached or shared across jobs. Remove the listener in cleanup, or use a one-off wait for a single expected event.
Promise rejection appears as unhandled A wait promise was created but not eventually awaited or caught, often because the action failed first. Keep the wait and action in one scoped operation, handle failures, and ensure every created promise is awaited or deliberately cancelled.

8. Performance, reliability, and cost

  • Match narrowly. A specific response predicate reduces accidental matches and helps the scraper proceed as soon as the needed data arrives.
  • Wait only for necessary conditions. Waiting for a complete page load when extraction depends on one API response adds avoidable latency. Conversely, extracting immediately after a response may be too early if the page must render first.
  • Set bounded timeouts. Per-operation limits prevent one stalled page from occupying a worker indefinitely. Choose values based on the site and your service’s overall deadline; the right value is workload-specific.
  • Keep listeners scoped. Remove persistent handlers after use and close pages and browsers in cleanup paths. This keeps long-running scraping processes easier to reason about.
  • Log conditions, not just stack traces. Record the page URL, expected event or selector, elapsed time, and status where available. This distinguishes a slow site from a wrong predicate.
  • Account for browser operations. A browser-based scraper consumes time and compute while it launches, loads, executes JavaScript, and waits. Reuse browser processes where appropriate in a managed worker, while keeping page and context state isolated according to the job.
  • Handle partial outcomes. A matching HTTP response can still be an error status; a URL transition may have no document response; an element may exist but contain no useful data. Validate the result before accepting a scrape.

9. Or skip the browser setup

If the job is to capture a page as an image or PDF rather than interact with its data, ScreenshotNeo provides a one-call screenshot API and an MCP server for AI agents. A GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo website and API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. This is useful for screenshots and PDFs; it does not replace browser automation when your scraper must click through a flow or extract structured page data.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

10. FAQ

Should I use Playwright or Puppeteer?

Both provide browser automation and event waits. Choose based on your existing project, browser needs, and the API guidance for the package version you use. The synchronization principle—subscribe before the trigger—applies to both.

Does a successful response mean the page is ready?

Only if your extraction depends on that response. If you need content rendered into the DOM afterward, wait for the corresponding element or state too.

Can a navigation wait return no response?

Yes. In Puppeteer, a History API URL change or anchor navigation can resolve without a main-resource response. Check the URL or page state when that outcome matters.

Is a fixed delay ever useful?

It can help during debugging or when a deliberate delay is itself part of the task. For normal readiness, an observable response, URL, or element condition is more reliable.

Can I use ScreenshotNeo to scrape structured data?

ScreenshotNeo captures screenshots and PDFs and exposes page information through its MCP tools. Use browser automation when the task requires interacting with a page and extracting structured data from it.

Sources