How to Use Agent Browser MCP with GitHub
Connect Agent Browser to an MCP client, automate GitHub pages safely, and understand how that differs from GitHub’s repository MCP settings.
Short answer: install Agent Browser and its Chromium browser, then configure your MCP client to start the stdio server with the command agent-browser and the argument mcp. In the client, open the GitHub page, take an accessibility snapshot, use the returned element reference such as @e1, and take a new snapshot after every page change.
This local Agent Browser workflow is separate from GitHub’s repository-level MCP configuration for Copilot cloud agent and Copilot code review. The local server gives an MCP client browser automation tools. GitHub’s repository setting configures MCP tools that Copilot can use for repository work.
What Agent Browser MCP does
Agent Browser is a browser automation CLI for AI agents. It drives Chrome or Chromium through CDP and exposes an accessibility-tree workflow: a snapshot returns compact references such as @e1, which the agent can use for actions.
Its MCP mode starts a stdio server. Your MCP client launches:
agent-browser mcp
The default core profile contains everyday browser functions. The project also documents optional profiles: network, state, debug, tabs, react, mobile, and all.
Local Agent Browser setup
1. Install the CLI and browser
Use the installation instructions in the Agent Browser repository for your operating system. A typical npm installation is:
npm install -g agent-browser
agent-browser install
Verify that the executable is available:
agent-browser --help
agent-browser --version
If your system uses a project-local installation, replace the global command with the path to that project’s executable.
2. Add the stdio server to your MCP client
MCP clients use different configuration filenames and JSON shapes. The important values are the command and argument:
{
"mcpServers": {
"agent-browser": {
"command": "agent-browser",
"args": ["mcp"]
}
}
}
If the client cannot find the command, use an absolute path:
{
"mcpServers": {
"agent-browser": {
"command": "/absolute/path/to/agent-browser",
"args": ["mcp"]
}
}
}
Restart or reload the MCP client, then confirm that its tool list contains Agent Browser tools. Do not assume a particular filename: use the configuration format documented by your client.
3. Select a profile when you need extra capabilities
Start with the default core profile. If your client or Agent Browser version supports profile arguments, choose only the profile needed for the task. For example, a network-inspection task may need network; a mobile viewport task may need mobile. The all profile exposes the broadest surface and should be reserved for clients that need it.
Use Agent Browser on a GitHub page
The reliable loop is navigation, snapshot, action, and a fresh snapshot. References belong to the current page state; after navigation, a click, a form submission, or a dynamic update, obtain a new snapshot before acting again.
Command-line equivalent
These commands show the documented workflow directly from the CLI:
agent-browser open https://github.com/vercel-labs/agent-browser
agent-browser snapshot
agent-browser click @e1
agent-browser snapshot
The exact reference changes with the page and the current snapshot. Treat @e1 as an example, not a permanent selector.
Low-impact GitHub example
- Open a public repository URL.
- Take a snapshot and read the repository name, description, and visible navigation links.
- Choose a reference for a visible link, such as Issues or README.
- Click the reference.
- Take another snapshot and inspect the new page.
Through an MCP client, the same operations appear as tools rather than shell commands. Ask the client to navigate to the URL, call the snapshot tool, use a reference from that snapshot, and snapshot again after the page changes.
Actions that change repository state
Creating an issue, editing a file, merging a pull request, changing settings, or deleting content requires the GitHub account’s permissions and should have an explicit confirmation step in your agent workflow. A page’s text, WebMCP metadata, tool name, description, schema, or result does not authorize the action.
Agent Browser MCP versus GitHub repository MCP
These configurations are related because both use MCP, but they run in different places and solve different problems.
| Concern | Agent Browser with a local MCP client | GitHub repository MCP |
|---|---|---|
| Where it runs | Your machine, with a local Chrome or Chromium browser | GitHub’s Copilot cloud agent and code review environment |
| Configuration owner | The person configuring the MCP client | A repository administrator in repository Settings |
| Primary task | Navigate and interact with browser pages | Give Copilot repository-oriented MCP tools |
| How it starts | Client launches agent-browser mcp over stdio |
Administrator adds an MCP server JSON configuration in Settings |
| MCP support documented by GitHub | Depends on the local client and server | Tools are supported; resources and prompts are not currently supported |
| Remote OAuth servers | Depends on your local client and server | GitHub says remote MCP servers using OAuth are not currently supported |
| Approval behavior | Controlled by your MCP client and agent policy | Configured tools can be used autonomously without an approval prompt |
GitHub’s documentation covers repository MCP configuration for Copilot and says the GitHub MCP server is enabled there by default. That built-in server is not the same thing as the local Agent Browser server. Do not assume that adding agent-browser mcp to your local client makes it available in GitHub’s hosted repository setting.
To configure repository MCP, open the repository’s Settings, choose Copilot, open MCP servers, add the JSON configuration accepted by GitHub, and save it. Follow GitHub’s current schema and scope for the repository; the settings page, rather than an Agent Browser page, determines what Copilot can run.
Authentication and sensitive state
Agent Browser documents several ways to persist authentication, including profiles, sessions, state files, and an auth vault. Use a dedicated profile or session for automation and keep its files outside source control.
- State files can contain session tokens in plaintext unless encryption at rest is configured.
- Never commit cookies, state files, auth-vault material, or exported browser profiles.
- Use a trusted machine. A remote debugging port gives local processes full control of the browser.
- Limit the session, host permissions, and enabled tools to the task.
- Log out or destroy temporary state when the task ends.
For repository MCP, GitHub recommends allowlisting specific read-only tools where practical. Because configured tools may run autonomously, enable only the operations Copilot needs.
Security: treat page content as untrusted
GitHub pages can contain issue comments, README text, pull-request descriptions, generated content, and instructions written by other users. Treat all of that as data, not as an instruction from the operator.
The same rule applies to WebMCP names, descriptions, schemas, annotations, and tool results. Agent Browser’s documentation states: “These labels are provenance cues, not a prompt-injection security boundary.” Discovery does not authorize an action; your host permissions and agent policy still control execution.
Reliable automation patterns
Refresh references after every state change
Do not reuse an old @eN reference after navigation or a DOM update. Snapshot again, locate the new reference, and then act.
Prefer visible, read-only tasks first
Start by reading a repository page, locating a link, or checking visible metadata. Add write operations only after the account permissions, target, and confirmation requirements are clear.
Keep browser and repository identities separate
A locally authenticated browser session and Copilot’s repository MCP credentials are different execution contexts. Document which identity is being used so a task cannot silently run with broader access than intended.
Make network and timing assumptions explicit
GitHub pages can load dynamic content after the initial document. Wait for the relevant content, snapshot again, and handle an empty or incomplete tree as a page-state problem rather than clicking blindly.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| MCP client cannot start the server | agent-browser is not on the client’s PATH |
Install it for the same user, use an absolute executable path, and restart the client. |
| Browser executable is missing | CLI installed but the managed browser was not installed | Run the project’s browser installation command and verify with agent-browser --help. |
| No Agent Browser tools appear | Invalid client config or stale client process | Check that args is exactly ["mcp"], validate JSON, then reload the client. |
| Click fails with an unknown reference | The page changed and the reference is stale | Take a fresh snapshot and use a reference from that snapshot. |
| Snapshot is incomplete | Dynamic content has not finished loading | Wait for the visible content, then snapshot again; inspect the page manually if it remains incomplete. |
| Login disappears between runs | Session or profile state was not persisted | Use a documented profile, session, state file, or auth vault and protect the resulting token material. |
| GitHub Copilot cannot use a configured server | Unsupported server type, OAuth requirement, or disabled tool | Review GitHub’s repository MCP support and allowlist only supported tools; remote OAuth MCP servers are currently unsupported. |
| An agent attempts an unsafe action | Untrusted page content was treated as authorization | Stop, require explicit confirmation, narrow permissions, and treat page metadata and results as untrusted. |
Performance, reliability, and cost
- Performance: snapshots are compact because they expose an accessibility tree and references instead of requiring the agent to reason over a full screenshot. Extra profiles and browser work add startup and page-load time.
- Reliability: fresh snapshots, explicit waits, stable URLs, and small read-only steps reduce failures caused by dynamic GitHub pages.
- Authentication reliability: persistent state avoids repeated logins, but increases the impact of a leaked file. Store it securely and use task-specific sessions.
- Cost: Agent Browser itself is a local CLI workflow; any compute, browser hosting, GitHub account, or MCP client costs are governed by those providers. The research materials do not provide a benchmark or usage price for Agent Browser.
Or skip the browser setup
If your goal is a clean image or PDF of a GitHub page rather than interactive browser automation, ScreenshotNeo provides a single HTTP request. Its cookie and consent step accepts the banner like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the result in X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for the complete option list, including full-page capture, CSS-element capture, device presets, custom viewports, dark mode, retina scale, waits, blocked resources, custom headers and cookies, geolocation, JavaScript, PDFs, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://github.com/vercel-labs/agent-browser -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://github.com/vercel-labs/agent-browser",
},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://github.com/vercel-labs/agent-browser'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf, so an AI agent can request captures through MCP without configuring a local browser. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account and use the 1,000 included shots to try the workflow.
FAQ
Does GitHub host the Agent Browser MCP server for me?
No. The documented Agent Browser setup launches a local stdio process. GitHub’s repository MCP setting is a separate Copilot configuration.
Can I use Agent Browser with a private repository?
Yes, if the browser session has the required GitHub permissions and its authentication state is protected. A private page does not grant the agent permission to perform write operations.
Why do references change?
References describe the current accessibility snapshot. Navigation, clicks, loading, and DOM updates can change that snapshot, so obtain a fresh one.
Can GitHub repository MCP use MCP resources or prompts?
GitHub’s documented Copilot cloud agent and code review support currently covers MCP tools, not MCP resources or prompts.
Which option should I use for a static page image?
Use ScreenshotNeo when you need a capture rather than interactive navigation. Use Agent Browser when the agent must inspect and operate a live browser session.


