How AI Browser Automation Can Work Without Playwright
Use CDP, Chrome DevTools MCP, WebDriver BiDi, Puppeteer, or an agent runtime to automate browsers without Playwright.
Yes. Playwright is optional. An AI browser agent can control Chrome through the Chrome DevTools Protocol (CDP), connect to Chrome DevTools MCP, use the standards-based WebDriver BiDi protocol through Selenium, or use Puppeteer, which supports CDP and WebDriver BiDi. Agent runtimes such as Browser Use can reuse a local Chrome profile or connect to a hosted browser over CDP.
Choose the route by capability: Chrome DevTools MCP is the shortest path to an agent operating a live Chrome session; CDP gives Chromium-specific control; WebDriver BiDi gives a cross-browser, event-driven standard; Puppeteer is a practical JavaScript driver; an agent runtime supplies planning and tool orchestration.
What changes when Playwright is removed
Playwright normally combines browser launch, selectors, waits, screenshots, tracing, and multi-browser support behind one API. Without it, you assemble those pieces yourself.
| Route | Best fit | Trade-off |
|---|---|---|
| Chrome DevTools MCP | An AI agent operating a live Chrome | Chrome-focused; the attached profile can expose sensitive sessions |
| Direct CDP | Low-level Chromium control, network, and performance diagnostics | Vendor-specific rather than cross-browser |
| WebDriver BiDi with Selenium | Cross-browser automation with streaming events | More protocol and driver setup |
| Puppeteer | JavaScript teams and existing Puppeteer code | Still requires browser lifecycle and session design |
| Browser Use or another agent runtime | Task planning and managed browser sessions | Inspect the underlying protocol before assuming portability or anti-bot behavior |
1. Connect an AI agent with Chrome DevTools MCP
Chrome DevTools for agents documents an MCP server that connects an AI agent to a live browser instance. The agent can inspect the DOM, evaluate JavaScript, capture screenshots, inspect network and performance data, and act in an already-open session.
Start a dedicated Chrome session
google-chrome --remote-debugging-port=9222 --user-data-dir=/tmp/agent-chrome
Use a dedicated profile. An attached agent can read page content and authenticated tabs, including cookies and local storage. Do not attach your daily profile or a high-privilege account.
Run the MCP server
npx chrome-devtools-mcp --browserUrl http://127.0.0.1:9222
Configure your MCP client with that command, then expose the server tools to the model. Keep the browser and MCP processes on the same isolated machine or container.
Useful agent loop
- Ask the agent to inspect the current page and identify the target element.
- Require a read-only check before any mutation.
- Perform one action, then verify the URL, DOM state, and console errors.
- Capture a screenshot or page state as an audit artifact.
2. Drive Chromium directly with CDP
CDP is Chrome and Chromium’s native debugging protocol. It fits workflows needing precise control over targets, JavaScript, screenshots, network interception, or performance events.
Runnable Node.js CDP screenshot
npm install chrome-remote-interface
const CDP = require('chrome-remote-interface');
(async () => {
const client = await CDP({ host: '127.0.0.1', port: 9222 });
const { Page, Runtime } = client;
await Page.enable();
await Runtime.enable();
await Page.navigate({ url: 'https://example.com' });
await Page.loadEventFired();
const shot = await Page.captureScreenshot({ format: 'png', captureBeyondViewport: true });
require('fs').writeFileSync('example.png', Buffer.from(shot.data, 'base64'));
await client.close();
})().catch(err => { console.error(err); process.exit(1); });
For an AI agent, expose narrow tools such as navigate, find_text, click, evaluate, and screenshot. Validate URLs and selectors, cap script execution time, and return structured results.
CDP details that matter
- Targets: pages, iframes, workers, and extensions are separate targets.
- Readiness:
loaddoes not mean application data is ready; wait for a selector or application signal. - Network: enable interception only when needed and disable it afterward.
- Headless mode: use headless Chrome for background jobs and visible mode while debugging.
- Isolation: use a fresh user-data directory per tenant or job when credentials must not mix.
3. Use WebDriver BiDi through Selenium
WebDriver BiDi is the W3C standard bidirectional protocol for browser automation. It adds a WebSocket connection so clients can receive events such as network requests, console messages, and JavaScript errors. It is the standards-first route when Firefox, Chrome, and other browsers matter.
Python Selenium example
python -m pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--no-sandbox')
driver = webdriver.Chrome(options=options)
try:
driver.get('https://example.com')
heading = WebDriverWait(driver, 20).until(EC.visibility_of_element_located((By.TAG_NAME, 'h1')))
print(heading.text)
driver.save_screenshot('example.png')
finally:
driver.quit()
Selenium’s BiDi APIs are version-sensitive. Pin Selenium and browser versions together, then enable the BiDi modules documented for that release.
4. Use Puppeteer without Playwright
Puppeteer controls Chrome through CDP and can also select WebDriver BiDi. It is useful for JavaScript teams with existing Puppeteer helpers.
npm install puppeteer
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle2', timeout: 30000 });
await page.screenshot({ path: 'example.png', fullPage: true });
console.log(await page.title());
} finally { await browser.close(); }
})().catch(err => { console.error(err); process.exit(1); });
Verify browser and Puppeteer versions together when selecting WebDriver BiDi. CDP remains preferable for Chrome-only domains.
5. Use an agent-oriented runtime
Browser Use documents reusing a local Chrome profile and connecting to hosted browsers through CDP. This supplies a task-oriented interface instead of exposing every protocol command.
python -m pip install browser-use
from browser_use import Agent
agent = Agent(task='Open https://example.com and report the main heading')
agent.run()
Check how the runtime launches browsers, stores credentials, handles retries, and selects protocols before assuming browser coverage or anti-bot behavior.
How to choose a route
| Requirement | Recommended route |
|---|---|
| Operate an existing logged-in Chrome tab | Chrome DevTools MCP |
| Chromium-only debugging or interception | Direct CDP |
| Cross-browser portability and events | WebDriver BiDi with Selenium |
| JavaScript codebase with Puppeteer skills | Puppeteer |
| Natural-language tasks with managed sessions | An agent runtime such as Browser Use |
Reliability, security, and performance checklist
- Use a dedicated profile, container, or VM; never share personal cookies with an agent.
- Grant least-privilege accounts and require confirmation for irreversible actions.
- Set navigation, action, and total-task timeouts. Retry only idempotent steps.
- Prefer stable selectors and semantic checks over coordinates.
- Keep screenshots and event logs out of the model prompt unless needed.
- Limit concurrency per browser profile.
- Record protocol, browser, driver, and library versions.
- Capture only the needed element or viewport when possible.
Common errors and fixes
| Error | Cause | Fix |
|---|---|---|
| Cannot connect to CDP | Chrome lacks remote debugging or the port is blocked | Launch an isolated profile with the debugging port and test locally. |
| Wrong tab | Multiple targets are open | List targets, select by URL or title, and reject unexpected origins. |
| Element not found | SPA content or iframe is not ready | Wait for a semantic selector, inspect frames, and check console errors. |
| Blank screenshot | Capture occurred before paint, navigation failed, or a bot check appeared | Wait for a visible element and verify the URL and response. |
| Missing BiDi events | Browser, driver, or Selenium version lacks the event | Pin compatible versions and enable the release’s BiDi module. |
| Session data leaked | Shared profile or reused cookies | Use a new profile per tenant or job. |
| Repeated destructive actions | No idempotency or post-action verification | Add dry-run mode, confirmation gates, and state checks. |
Or skip the browser setup
ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the verdict and billing.
See the ScreenshotNeo API docs for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page and element capture, device presets, dark mode, custom CSS and JavaScript, waits, request blocking, headers and cookies, geolocation, resizing, caching, signed links, async webhooks, bulk capture, PDF, HTML/CSS rendering, and usage reporting. Its MCP tools are take_screenshot, get_page_info, and capture_pdf.
There are 1,000 free screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free. Create a free ScreenshotNeo account.
FAQ
Is CDP an AI protocol?
No. CDP is a browser debugging protocol. Your agent or MCP server turns model decisions into CDP commands.
Does WebDriver BiDi replace CDP everywhere?
No. BiDi improves portability; CDP still exposes Chromium-specific domains and diagnostics.
Can I reuse my logged-in browser?
Technically yes, but use a dedicated profile and least-privilege account because the agent can access session data.
Which option is cheapest?
Local CDP, Selenium, and Puppeteer use your own infrastructure. Hosted runtimes and screenshot APIs add service charges.


