How to Use Agent-Browser with Playwright
Learn how agent-browser relates to Playwright, install the CLI, automate pages with snapshots, and choose the right tool for coding-agent workflows.
Short answer: use agent-browser through its own CLI. Install the package, install a supported Chrome browser, open a page, take a snapshot, and interact with the references or selectors returned by the snapshot. You do not install Playwright separately just to use the documented agent-browser daemon. The default daemon uses Playwright internally, while an experimental native mode uses direct CDP/WebDriver instead.
This distinction matters because agent-browser and Playwright’s playwright-cli are separate command-line tools. The commands look similar, but they have different installation paths and runtime behavior.
What “with Playwright” means
There are three possible meanings:
- Using agent-browser: you run the agent-browser CLI. Its default Node.js daemon uses Playwright internally. The project README states that end users do not need Playwright or Node.js installed to run the daemon: see the official agent-browser repository and command reference.
- Using agent-browser native mode: the project describes an experimental pure-Rust daemon based on direct CDP/WebDriver. It has different browser and feature support, so check the current native-mode changelog before selecting it.
- Using Playwright’s coding-agent CLI: Microsoft documents
playwright-clias a separate tool with its own package, prerequisites and commands: Playwright coding-agent documentation.
If you want to write a Playwright test or use a Playwright Page object directly, the reviewed agent-browser documentation does not establish an API bridge for embedding agent-browser commands in that script. Use the documented CLI workflow unless you verify a separate integration.
Install agent-browser
Global installation
npm install -g agent-browser
agent-browser install
The second command downloads Chrome for Testing for the local workflow.
Project-local installation
npm install agent-browser
agent-browser install
Use the project-local form when you want the dependency and its version recorded in your application. Invoke the binary through your package manager or an npm script.
Linux dependencies
On Linux systems that need browser libraries, the project documents:
agent-browser install --with-deps
Other distribution routes
The project also documents Homebrew on macOS and Cargo distribution. Building from source has separate requirements, including Node.js 24+, pnpm 11+ and Rust. Those source-build requirements do not apply to ordinary CLI use.
Your first agent-browser session
Run this sequence in a terminal:
agent-browser open https://example.com
agent-browser snapshot
agent-browser click @e2
agent-browser snapshot
agent-browser close
The reference @e2 is an example. Your snapshot will produce its own references. Never assume that a reference from one page or state will have the same value in another session.
What each command does
| Command | Purpose |
|---|---|
open URL |
Starts or reuses a session and navigates to a URL. |
snapshot |
Prints a compact representation of the page with references for interactive elements. |
click REF |
Clicks an element identified by a snapshot reference. |
close |
Closes the browser session. |
Interact with snapshot references and selectors
Use a snapshot reference
agent-browser open https://example.com/login
agent-browser snapshot
agent-browser click @e4
agent-browser snapshot
References are convenient for an agent because they are generated from the current page state. After a click, navigation, modal opening or DOM update, take another snapshot before using more references. The project specifically recommends refreshing the snapshot after dismissing an element that covered a click target.
Use a CSS selector
agent-browser open https://example.com/login
agent-browser fill "#email" "dev@example.com"
agent-browser fill "#password" "correct-horse-battery-staple"
agent-browser click "button[type=submit]"
agent-browser snapshot
Selectors are useful when a stable attribute is known. Prefer specific IDs, labels or test attributes over fragile positional selectors.
Find an element by role
agent-browser open https://example.com
agent-browser find role button click --name "Submit"
agent-browser snapshot
Role-based commands make intent clearer and can survive some markup changes better than a long CSS path.
Common commands for a complete workflow
Fill a form and read text
agent-browser open https://example.com/contact
agent-browser fill "input[name=name]" "Ada Lovelace"
agent-browser fill "input[name=email]" "ada@example.com"
agent-browser fill "textarea[name=message]" "Please send the documentation."
agent-browser click "button[type=submit]"
agent-browser snapshot
agent-browser get text "body"
Capture a screenshot
agent-browser open https://example.com
agent-browser screenshot page.png
Use a screenshot when you need visual evidence. Use snapshot when the agent needs structured page state and actionable references.
Work with multiple tabs
agent-browser open https://example.com
agent-browser tab new https://example.org
agent-browser tab list
agent-browser tab switch 0
agent-browser snapshot
Check the current command reference for exact tab subcommands and flags because the interface can change between releases.
Attach to an existing browser over CDP
agent-browser connect <cdp-endpoint>
agent-browser snapshot
This is useful when the browser runs in another process or on remote infrastructure. The endpoint, authentication and network path are deployment-specific.
Snapshot-driven agent pattern
A reliable agent loop is:
- Navigate to the target URL.
- Take a snapshot.
- Choose an action from the current references, roles or selectors.
- Perform one action.
- Take a fresh snapshot after any state-changing action.
- Repeat until the task is complete.
- Capture text or a screenshot as evidence.
- Close the session.
Do not cache references across navigation. A reference can become invalid when the page rerenders, a dialog closes, a route changes or a lazy component appears.
Agent-browser versus Playwright CLI
| Question | agent-browser | Playwright CLI |
|---|---|---|
| Primary interface | agent-browser commands |
playwright-cli commands |
| Typical use | Agent-controlled browsing with compact snapshots and references | Playwright’s coding-agent workflow |
| Separate installation? | Install agent-browser; its default daemon uses Playwright internally | Install the Playwright CLI package as documented by Microsoft |
| Direct Playwright test API | Not established by the reviewed documentation | Designed for the Playwright toolchain |
| Prerequisites | Follow agent-browser’s current installation guide | Playwright documents Node.js 20+ for its CLI |
Do not copy a playwright-cli command into an agent-browser script or assume that an agent-browser reference can be passed to a Playwright Page object.
Native mode and browser compatibility
The agent-browser changelog describes an experimental native Rust daemon that communicates through CDP/WebDriver. At the cited March 3, 2026 entry, native mode differs from the default daemon in several ways:
- Firefox and WebKit are not supported in native mode.
- Playwright trace format and HAR export are unavailable.
- Network routing uses CDP Fetch rather than Playwright’s route API.
- The project advises closing the session before switching modes.
These details are version-sensitive. Recheck the current changelog before using native mode in production.
Remote browser execution
Local Chrome is the simplest deployment. If your runtime cannot launch a browser, the project documents remote browser infrastructure such as Browserbase as an optional route. Remote execution adds credentials, network configuration and service-account management; it is not required for the basic local workflow.
Troubleshooting
“Command not found: agent-browser”
Cause: the package is not installed globally, or the global npm bin directory is not on PATH.
Fix: install globally with npm install -g agent-browser, or install locally and invoke the project-local binary through your package manager.
Browser executable is missing
Cause: Chrome for Testing has not been downloaded for the agent-browser session.
Fix: run agent-browser install. On Linux, use agent-browser install --with-deps when required system libraries are absent.
Linux launch fails with a shared-library error
Cause: the host image lacks browser libraries.
Fix: rerun the documented install command with --with-deps, or add the required packages to the container image used by your deployment.
A reference such as @e2 no longer works
Cause: the page changed after navigation, a click, a modal transition or a rerender.
Fix: call agent-browser snapshot again and use the new reference. Prefer a stable selector or role when the element has a predictable name.
A click is blocked by a modal or consent banner
Cause: an overlay is covering the target.
Fix: inspect the snapshot, dismiss the overlay, then take a fresh snapshot before clicking the original target.
The command behaves differently after changing runtime mode
Cause: native mode and the default Playwright-based daemon do not expose identical capabilities.
Fix: close the current session, switch modes deliberately, and check the current changelog for unsupported features.
You are trying to pass agent-browser into a Playwright test
Cause: the CLI and Playwright’s programmatic API are being mixed.
Fix: either automate through the agent-browser CLI, or write a normal Playwright script. The reviewed sources do not document an embedding adapter between them.
Performance and reliability checklist
- Reuse a browser session for related actions instead of launching a new process for every command.
- Take snapshots after state changes, but avoid unnecessary snapshots when the page has not changed.
- Use stable IDs, labels, roles or test attributes instead of deeply nested CSS selectors.
- Record the URL and final snapshot or screenshot when an agent action must be audited.
- Pin the package version in project-local deployments and rerun browser installation after browser or Playwright upgrades when the documentation requires it.
- Close sessions in cleanup handlers so failed jobs do not leave browser processes running.
- Use a remote browser only when local execution is unsuitable; account for network latency and credential setup.
Or skip the browser setup
If your goal is a clean image of a web page rather than interactive browser control, ScreenshotNeo provides a one-request screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before capture; bot checks, blank pages, failed loads, timeouts and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the complete ScreenshotNeo API documentation for options such as full-page capture, CSS-selector element capture, device presets, dark mode, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, PDF settings, caching, signed links, async jobs and bulk capture.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = await res.arrayBuffer();
// Save bytes with your runtime's file API.
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Do I need to install Playwright before agent-browser?
No. The documented agent-browser CLI handles its default daemon setup and uses Playwright internally.
Is agent-browser a Playwright plugin?
The reviewed documentation presents it as a standalone CLI, not as a Playwright Page-object plugin.
When should I use playwright-cli instead?
Use playwright-cli when you specifically want Playwright’s coding-agent command workflow and its documented setup.
Why must I refresh snapshots?
References describe the current page state. Navigation, rerenders and dialogs can invalidate them.
Can I use agent-browser without a local browser?
Yes, through a documented remote browser deployment, but the basic workflow assumes a locally available supported browser.


