ScreenshotNeo

BlogAI agents

How to Use Midscene.js with MCP for Browser Automation

Configure Midscene.js as an MCP browser tool, automate a page with an AI assistant, verify results, and choose between MCP, Bridge, CDP, and headless modes.

By the ScreenshotNeo team1 October 20267 min read

Direct answer: configure Midscene’s MCP server in your AI client, provide a supported multimodal model and its credentials, then ask the assistant to call Midscene tools such as midscene_navigate, midscene_aiTap, midscene_aiInput, midscene_aiWaitFor, midscene_aiAssert, and midscene_screenshot. MCP gives the assistant a browser-control interface; it is separate from Midscene’s CLI, Chrome Bridge, CDP, Playwright, and Puppeteer integration paths.

What you will build

The finished setup lets an MCP-compatible client ask Midscene to open a page, inspect tabs, interact with visible elements using natural language, wait for a condition, verify the result, and capture a screenshot. Midscene’s MCP documentation also lists a tool that returns Playwright example code.

Prerequisites

  • Node.js and npm, so your MCP client can launch npx.
  • An MCP-compatible AI client, such as a client that supports custom local MCP servers.
  • A supported multimodal model provider and credentials. The provider key is not bundled with Midscene.
  • For the Chrome route: the Midscene Chrome extension, Bridge Mode enabled, and permission for the connection.

Model environment variable names differ by provider. Use Midscene’s current model-selection documentation for the provider you chose.

Configure the Midscene MCP server

Add a server entry to your AI client’s MCP configuration. The following mirrors the official example and uses placeholders where your provider-specific values belong:

{
  "mcpServers": {
    "mcp-midscene": {
      "command": "npx",
      "args": ["-y", "@midscene/mcp"],
      "env": {
        "MIDSCENE_MODEL_NAME": "YOUR_SUPPORTED_MODEL",
        "OPENAI_API_KEY": "YOUR_PROVIDER_KEY",
        "MCP_SERVER_REQUEST_TIMEOUT": "800000"
      }
    }
  }
}

Adjust before running:

  1. Replace YOUR_SUPPORTED_MODEL with a model supported by your selected provider.
  2. Set the environment variable required by that provider. The example uses OPENAI_API_KEY; another provider may require different names.
  3. Keep the timeout as a starting point or change it for long browser tasks. It is expressed as a string in the JSON environment object.
  4. Restart or reload the MCP client after saving its configuration.

Keep credentials in the client’s environment or secret manager. Do not commit them to a repository or paste them into prompts.

Run a browser task through MCP

Tool names can vary slightly by client presentation, but the MCP guide lists these Midscene tools:

Tool Purpose
midscene_navigate Navigate the current tab to a URL.
midscene_get_tabs List open tabs and their IDs.
midscene_set_active_tab Select a tab by ID.
midscene_aiTap Click an element described in natural language.
midscene_aiInput Enter text into a described field.
midscene_aiHover Hover over a described element.
midscene_aiKeyboardPress Press a keyboard key or combination.
midscene_aiScroll Scroll to find or reveal content.
midscene_aiWaitFor Wait for a visible page condition.
midscene_aiAssert Check that a page condition is true.
midscene_screenshot Capture the current page.
midscene_playwright_example Return a Playwright example for the task.

Illustrative task prompt

This is a writer-created workflow example, not a claim that it was executed:

Open https://example.com.
1. Inspect the open tabs and select the active page if necessary.
2. Find the page search field and enter "browser automation".
3. Submit the search.
4. Wait until visible search results appear.
5. Assert that the results heading is visible.
6. Capture a screenshot and report what you observed.

A staged request is easier to diagnose than one large instruction. Ask the assistant to show the tool result after each important step, especially after navigation, tab selection, waits, and assertions.

Reliable task sequence

  1. Navigate: provide an explicit HTTPS URL and avoid starting with a destructive action.
  2. Inspect tabs: call midscene_get_tabs when the client may already have several tabs.
  3. Select the target: use midscene_set_active_tab with the intended tab ID.
  4. Interact: describe one element and one action at a time with midscene_aiTap, midscene_aiInput, hover, scroll, or keyboard tools.
  5. Wait: use midscene_aiWaitFor for a visible, meaningful condition such as a results heading or confirmation message.
  6. Assert: call midscene_aiAssert and state the exact expected outcome.
  7. Capture evidence: use midscene_screenshot after the assertion. Screenshots document the visible state; they do not replace an assertion.

Choosing a Midscene browser mode

MCP is the AI-client tool interface. Midscene also offers browser execution modes and application integrations. Choose them independently:

Path Browser location Best for What you configure
MCP server Controlled by the Midscene MCP process and client setup Letting an AI assistant call browser tools MCP JSON, model provider, credentials, timeout
Default CLI Headless Puppeteer browser Repeatable automation without a desktop session CLI command and task inputs
--bridge Your existing desktop Chrome Tasks requiring the logged-in desktop session or visible browser Extension, Bridge Mode, connection permission
--cdp <ws-endpoint> An existing browser exposed through Chrome DevTools Protocol Remote or infrastructure-managed browsers Reachable CDP WebSocket endpoint
Playwright/Puppeteer SDK Browser selected by your application Code-owned workflows, CI, and generated scripts JavaScript/TypeScript application and browser lifecycle

Midscene’s broader project supports direct scripts, CLI modes, Playwright and Puppeteer integrations, and a Chrome extension Bridge route. Those alternatives are not additional MCP configuration steps.

Chrome Bridge setup and limits

For Bridge mode, install the Midscene Chrome extension, switch it to Bridge Mode, allow the connection, and configure a supported model. Bridge mode follows the active desktop Chrome settings. The Bridge documentation says it ignores these automation options:

  • userAgent
  • viewportWidth and viewportHeight
  • deviceScaleFactor
  • waitForNetworkIdle
  • cookie
  • extraHTTPHeaders
  • downloadPath
  • chromeArgs

Set those values in the desktop browser when Bridge mode requires them. A remote Bridge connection binds to 127.0.0.1 by default. If you bind it to another interface, restrict access to a trusted network and firewall it; do not expose the bridge server on a public network.

Bridge scripting example

This illustrates Midscene’s separate Bridge scripting API, not MCP client configuration:

import { AgentOverChromeBridge } from '@midscene/web/bridge-mode';

const agent = new AgentOverChromeBridge();
await agent.connectNewTab('https://example.com');
await agent.aiAction('Click the page link labeled More information');
await agent.aiAssert('The destination page is visible');
await agent.destroy();

The documented Bridge example installs @midscene/web and tsx, then runs the script with tsx. Adapt the selector and assertions to your page.

Performance, reliability, and cost considerations

  • Latency: visual interpretation, page loading, waits, and model requests all add time. Use a specific wait condition instead of an unnecessarily long fixed delay.
  • Reliability: begin with one harmless page action, assert the visible result, and save a screenshot. Break multi-page jobs into checkpoints so a failed step is identifiable.
  • Browser state: Bridge inherits desktop tabs, cookies, viewport, extensions, and permissions. Headless and CDP workflows make browser ownership more explicit.
  • Timeouts: increase MCP_SERVER_REQUEST_TIMEOUT only when a legitimate workflow needs it; a larger timeout can make failures take longer to surface.
  • Model cost: Midscene’s MCP setup requires a supported provider and its credentials. Provider pricing and quotas are separate from Midscene.
  • Security: treat browser control as access to every page and session available to that browser. Use a dedicated profile for sensitive automation and restrict remote endpoints.

Troubleshooting

Symptom Likely cause Fix
The MCP server will not start Node.js/npm is missing, or npx cannot download the package Check Node.js and npm, run npx -y @midscene/mcp manually, and inspect the client log.
Model or provider error Unsupported model name or wrong environment variable Follow the current provider-specific model documentation and set the matching credential variable.
Requests time out Slow page, model response, or timeout too low Use a narrower task, add a condition-based wait, then adjust MCP_SERVER_REQUEST_TIMEOUT if required.
The assistant acts in the wrong tab Several tabs are open Call midscene_get_tabs, then midscene_set_active_tab with the intended ID.
An element cannot be found It is below the fold, still loading, hidden, or described ambiguously Scroll, wait for a visible condition, and describe the element by its visible role or label.
An assertion fails after a successful action The page has not finished updating, or the expected text is wrong Wait for the actual post-action condition and assert a stable heading, status, or confirmation.
Bridge ignores a setting The option is unsupported in Bridge mode Configure the active desktop Chrome; Bridge ignores user agent, viewport, device scale, network-idle, cookies, headers, download path, and Chrome arguments.
Remote Bridge cannot connect Binding or firewall prevents access Use the correct interface and port, keep access on a trusted network, and add restrictive firewall rules.
A screenshot shows the wrong state Capture happened before navigation or rendering completed Wait for a visible condition and capture after the assertion.

Or skip the browser setup

If your goal is a clean screenshot rather than interactive browser control, ScreenshotNeo returns an image or PDF from one request. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Validation checklist

  • Confirm the MCP client sees mcp-midscene.
  • Verify the model name and provider credential.
  • Navigate to a low-risk test page.
  • Inspect and select the intended tab.
  • Perform one reversible interaction.
  • Wait for a visible result.
  • Assert the result.
  • Capture a screenshot and review the tool output.
  • If using Bridge, confirm extension permission and network restrictions.

FAQ

Does MCP replace Playwright or Puppeteer?

No. MCP exposes Midscene capabilities to an AI client. Midscene also supports direct Playwright and Puppeteer integrations for code-owned automation.

Do I need Chrome Bridge to use Midscene MCP?

No. Bridge is one browser route. The MCP setup and the CLI’s headless and CDP modes are separate choices.

Can MCP tasks generate reusable code?

Yes. The MCP tool list includes midscene_playwright_example, which can return a Playwright example for a task.

Is the provider API key included?

No. You must supply credentials for a supported model provider and follow that provider’s current variable names.

Should I expose a remote Bridge server publicly?

No. Keep remote access on a trusted network and use firewall restrictions.