ScreenshotNeo

BlogHow-to

How to Use Web APIs for Browser Automation

Learn how browser automation APIs, CDP, WebDriver BiDi, Puppeteer, Selenium, and Playwright fit together, with runnable setup and CI examples.

By the ScreenshotNeo team1 October 20269 min read

Short answer: browser automation APIs let a program launch or attach to a browser, navigate to pages, interact with controls, and observe browser events. In practice, you choose a library such as Puppeteer, Selenium, or Playwright, then use a browser-control protocol such as Chrome DevTools Protocol (CDP) or WebDriver BiDi underneath.

This guide explains the protocols, shows complete examples, and covers browser version pinning, CI, events, reliability, performance, costs, and common failures. Use automation only on sites and accounts where you have permission, and follow the site’s rules.

1. What “web APIs” means in browser automation

The phrase can mean two different things:

  • Page JavaScript APIs: APIs exposed inside a webpage, such as fetch, the DOM, and storage.
  • Browser automation APIs: APIs used by an external program to control a browser process. They can open pages, click, type, inspect network traffic, capture logs, and collect screenshots.

This article focuses on the second meaning. A typical flow is:

  1. Install or select a browser binary.
  2. Launch it or connect to an existing browser.
  3. Create a page or browsing context.
  4. Navigate and wait for the state your test needs.
  5. Interact with elements and observe events.
  6. Assert results, save artifacts, and close the browser.

2. CDP and WebDriver BiDi

Protocol What it provides Best fit Important limitation
CDP Commands and events for Chromium, Chrome, and other Blink-based browsers. Chrome-specific debugging, network and performance instrumentation, and tools built around Chromium. The tip-of-tree definitions change frequently and have no guaranteed backward compatibility. Pin compatible browser and library versions. Chrome DevTools Protocol documentation
WebDriver BiDi A W3C bidirectional WebSocket protocol. Automation code can receive events such as network requests, console messages, and JavaScript errors. Standards-oriented automation and event-driven workflows across supported browsers. Feature support is still being implemented by browsers and drivers. Selenium describes CDP support as temporary while BiDi implementations mature. Selenium WebDriver BiDi documentation

Classic WebDriver commands are request/response oriented. BiDi adds a long-lived two-way event stream. CDP remains useful when you need Chromium-specific capabilities, but its unstable protocol surface makes a library’s supported API safer for application code.

3. Choose a browser automation library

Library Languages Browser coverage Protocol and infrastructure strengths
Puppeteer JavaScript/Node.js Chrome and Firefox Maintained by Chrome’s Browser Automation team. Chrome uses CDP by default; Firefox uses BiDi by default, and production-ready BiDi support is available for both. Each release is tied to a browser release to protect compatibility. Puppeteer FAQ
Selenium Java, Python, C#, Ruby, JavaScript and more Broad browser and driver ecosystem More language bindings and Selenium Grid orchestration. Enable BiDi with the webSocketUrl capability. Selenium documentation
Playwright JavaScript/TypeScript, Python, Java, .NET Chromium, Firefox, WebKit Its own protocol gives the highest fidelity. connectOverCDP is Chromium-only and significantly lower fidelity than a Playwright protocol connection. Playwright BrowserType API

Choose using these axes: browser engines required, preferred language, need for a standards-based event stream, framework-specific features, grid or distributed execution, version alignment, and whether the browser runs visibly during development or headlessly in CI.

4. A reproducible Chrome setup

Chrome for Testing is a Chrome distribution intended for testing and automation. Its versioned downloads and matching ChromeDriver binaries let teams pin the browser and driver used in CI. Modern headless mode uses the same browser implementation as headful Chrome.

  1. Record the browser version in your project configuration.
  2. Install the matching driver when using WebDriver.
  3. Run headless in CI; run headful locally when diagnosing selectors or timing.
  4. Cache the pinned browser binary in CI, but invalidate the cache when the version changes.

Puppeteer can download a compatible Chrome for Testing binary by default and launches headless in its typical workflow. Selenium users should manage Chrome and ChromeDriver as a compatible pair.

5. Complete Puppeteer example (Node.js)

Install Puppeteer:

npm install puppeteer

Create automation.mjs:

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({
  headless: true
});

try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 30000 });

  const title = await page.title();
  const heading = await page.locator('h1').innerText();
  console.log({ title, heading });

  await page.screenshot({ path: 'example.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node automation.mjs. Prefer locators or stable attributes over brittle positional selectors. Add an explicit wait for the state that proves the action completed, such as a visible result or URL change.

Listening to browser events with Puppeteer

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  page.on('console', message => console.log('console:', message.type(), message.text()));
  page.on('pageerror', error => console.error('page error:', error.message));
  page.on('requestfailed', request => console.error('request failed:', request.url(), request.failure()?.errorText));
  await page.goto('https://example.com', { waitUntil: 'networkidle0' });
} finally {
  await browser.close();
}

6. Complete Selenium example with WebDriver BiDi (Python)

Install Selenium:

python -m pip install selenium

The following example enables the BiDi WebSocket capability and uses Selenium’s normal navigation API. Exact event APIs depend on the Selenium and browser versions you pin; consult the current BiDi documentation for the feature group you need.

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By

options = Options()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,900')
options.set_capability('webSocketUrl', True)

driver = webdriver.Chrome(options=options)
try:
    driver.get('https://example.com')
    heading = driver.find_element(By.CSS_SELECTOR, 'h1').text
    print({'title': driver.title, 'heading': heading})
    driver.save_screenshot('example.png')
finally:
    driver.quit()

For a distributed suite, Selenium Grid can route sessions to workers with the required browser and driver versions. Keep the same capability and timeout policy across workers.

7. Playwright across Chromium, Firefox, and WebKit

npm init -y
npm install playwright
npx playwright install
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
  await page.getByRole('heading', { name: 'Example Domain' }).waitFor();
  await page.screenshot({ path: 'example.png', fullPage: true });
} finally {
  await browser.close();
}

Use Playwright’s own connection when you need its full feature set. Its connectOverCDP path supports Chromium-based browsers only and is lower fidelity; launching an external browser with incompatible arguments can break features. See the BrowserType API.

8. Waiting, interaction, and state design

  • Wait for a meaningful condition, not an arbitrary long sleep: an element becoming visible, a URL matching, a response arriving, or a loading indicator disappearing.
  • Use separate timeouts for navigation, actions, and assertions so failures identify the slow stage.
  • Make tests independent: create a fresh context, seed required state, and clean up files and sessions.
  • Capture a screenshot, HTML, console log, and network failure details when a failure occurs.
  • Use stable test IDs or accessible roles where you control the application.

9. Running browser automation in CI

  1. Pin the automation library and browser version in lockfiles or a container image.
  2. Install the browser and matching driver during image creation, not during every test job.
  3. Run headless with a fixed viewport, timezone, locale, and deterministic test data.
  4. Set a global job timeout and a shorter per-step timeout.
  5. Upload screenshots, traces, logs, and HTML only when a test fails.
  6. Retry only known transient failures; do not hide selector or assertion bugs with unlimited retries.

Chrome’s official automation guidance covers Chrome for Testing, ChromeDriver, Puppeteer, headless mode, and CI workflows: Chrome automation documentation.

10. Troubleshooting common failures

Symptom Likely cause Fix
Browser fails to launch in CI Missing libraries, sandbox restrictions, or an incompatible binary. Use a maintained browser image, install the pinned Chrome for Testing build, and inspect the browser’s stderr. Avoid adding launch flags without understanding their security and compatibility impact.
“Session not created” Chrome and ChromeDriver versions do not match. Pin and install the matching pair; print both versions in CI logs.
Element not found The page has not reached the expected state, the selector changed, or content is inside a frame or shadow root. Wait for a state, verify the frame/context, prefer stable selectors, and save HTML on failure.
Click intercepted or element not interactable An overlay, animation, viewport issue, or disabled control blocks the action. Wait for visibility and enabled state, scroll into view, close the overlay through the UI, or use a larger deterministic viewport.
Navigation timeout Slow resources, never-ending requests, redirects, or a page that does not reach the chosen readiness event. Use an appropriate readiness condition, block unnecessary resources where safe, and log the URL and failed requests.
BiDi events are missing The browser, driver, Selenium version, or capability does not support that event group. Check the support matrix, set webSocketUrl, update compatible components together, and fall back to a documented library API.
CDP command breaks after an upgrade Tip-of-tree CDP definitions changed. Use the framework API where possible and pin compatible browser/protocol versions.
Automation behaves differently from a human session Different viewport, locale, storage, permissions, or application state. Make those inputs explicit and record them with each run. Automation does not grant permission to access a site or guarantee that bot checks will allow access.

11. Performance, reliability, and cost

Performance

  • Reuse a browser process when isolation allows it; create fresh contexts or pages for independent work.
  • Keep the browser close to the application under test to reduce network latency.
  • Use targeted waits and avoid sleeping for a fixed large interval.
  • Collect only the artifacts needed for every run; retain full traces for failures.
  • Parallelize at the worker level while respecting CPU, memory, and the site’s rate limits.

Reliability

  • Pin versions and update them deliberately.
  • Use deterministic data and reset state between tests.
  • Retry infrastructure failures with a cap, while surfacing product failures immediately.
  • Record browser, driver, library, OS, viewport, locale, and commit identifiers.

Cost

Self-hosted automation costs compute, storage for artifacts, and maintenance of browsers and drivers. Managed grids add service charges but reduce infrastructure work. Estimate concurrency from your longest test and required run window, then budget CPU and memory headroom for parallel browser processes.

12. Or skip the browser setup

If your goal is a clean screenshot rather than interactive browser testing, ScreenshotNeo provides a single HTTP request for PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page and element capture, device presets, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, async jobs, bulk capture, usage, and the OpenAPI specification.

curl -G 'https://api.screenshotneo.com/v1/shot' \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = new Uint8Array(await res.arrayBuffer());
await Bun.write('shot.webp', data);

An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

13. Frequently asked questions

How do I automate a browser with an API?

Install a library, launch or attach to a browser, navigate, interact with locators, wait for a verifiable state, collect results, and close the session. Puppeteer, Selenium, and Playwright provide these higher-level APIs.

What is the difference between CDP and WebDriver BiDi?

CDP is a Chromium-oriented command and event protocol with a rapidly changing tip-of-tree definition. WebDriver BiDi is a W3C bidirectional WebSocket protocol intended for standards-based cross-browser automation and event streaming.

Can Puppeteer automate Firefox?

Yes. Puppeteer supports Chrome and Firefox; its FAQ states that Chrome uses CDP by default, Firefox uses BiDi by default, and BiDi support is production-ready for both browsers.

Should I use Selenium, Playwright, or Puppeteer?

Use Selenium when language bindings or Grid orchestration matter, Playwright when you need Chromium, Firefox, and WebKit through one framework, and Puppeteer when a Node.js workflow centered on Chrome or Firefox fits your needs.

Can browser automation bypass bot detection?

No guarantee should be assumed. Sites can distinguish trusted and untrusted events, and Puppeteer documents trusted input behavior; automation does not grant access rights or defeat a site’s controls.