ScreenshotNeo

BlogAI agents

How to Capture a Website Screenshot with an AI Agent Using Selenium MCP

Set up a Selenium MCP server, ask an AI agent to capture a website, and troubleshoot browser, navigation, and screenshot issues.

By the ScreenshotNeo team4 October 202610 min read

To capture a website screenshot with an AI agent using Selenium MCP, connect an MCP-compatible client to a Selenium MCP server, then ask the agent to start a browser, navigate to a URL, capture a screenshot, and close the browser. The exact tool names, screenshot options, and where the image is returned or saved depend on the server you choose. This guide uses the documented @gaforov/selenium-mcp setup as an example; it is a community project, not an official Selenium component.

1. What Selenium MCP does

MCP gives an AI agent access to tools. A Selenium MCP server exposes browser actions through Selenium WebDriver, so the agent can operate a real browser without you writing every WebDriver command. A typical sequence is start_browser, navigate, optionally get_title, take_screenshot, then stop_browser. Tool names and arguments vary by server, so check the server’s current tool reference.

This approach is useful when you want to request a screenshot conversationally or inspect a live page. For repeatable test logic or a workflow that belongs in a test suite, a short Selenium script may be easier to review and maintain.

2. Requirements and server choice

The example below follows the gaforov/selenium-mcp README as reviewed on October 3, 2026. It documents Node.js 20 or later and an installed Chrome, Firefox, or Edge browser. It says Selenium Manager provisions a matching driver automatically. Confirm the current README before installing because package requirements, browser support, and tool interfaces can change.

Choice What to check
@gaforov/selenium-mcp Client setup, current tool names and arguments, screenshot output handling, supported browser, and maintenance. Its README reports 41 tools, including page snapshots, selector hints, batch execution, and tracing; these are project-reported capabilities.
mcp-selenium by angiejones Client setup and current browser requirements. Its README lists Chrome, Firefox, Edge, and Safari. It documents Safari-specific macOS setup, including enabling safaridriver and Remote Automation, and says Safari has no headless mode. Verify those details in its current README.
Direct Selenium script Choose this when you want browser operations and screenshot handling in code you own, without adding a third-party MCP tool server.

Selenium’s official documentation describes community MCP servers as third-party projects. Review the specific server’s maintenance, permissions, and behavior before connecting it to an agent.

3. Configure an MCP client

Claude Code

For the gaforov example, the README documents this command:

claude mcp add selenium -- npx -y @gaforov/selenium-mcp@latest

This registers an MCP server named selenium that Claude Code launches with npx. The @latest tag follows the documented example, but it can resolve to a newer package over time. For controlled environments, check whether the project documents a version-pinning option and use a version you have reviewed.

Claude Desktop, Cursor, or Windsurf

The same project documents an MCP server configuration with command npx and arguments -y and @gaforov/selenium-mcp@latest. MCP clients use different configuration locations and formats, so follow your client’s current instructions. The conceptual entry is:

{
  "mcpServers": {
    "selenium": {
      "command": "npx",
      "args": ["-y", "@gaforov/selenium-mcp@latest"]
    }
  }
}

Restart or reload the client as required, then confirm that the Selenium tools appear in its available tool list. Do not assume the same configuration file path works in every client.

4. Ask the agent to capture a screenshot

Once the server is connected, give the agent a concrete request. The gaforov README’s example is:

Use selenium-mcp to open Chrome, go to https://example.com, read the page title, take a screenshot, and close the browser.

The agent should use the tools exposed by the configured server to start the browser, navigate, verify the page title if available, take the screenshot, and stop the browser. Check the tool’s response to learn whether the image is attached to the conversation, returned as image content, or written to a file. The output behavior is server-specific.

Make the request more reliable

  • Name the browser if the server supports a browser parameter.
  • Ask the agent to report the page title or current URL before capture so navigation failures are visible.
  • Specify a wait condition when the site needs time to render, such as waiting for a known selector. Use a locator verified against the live page.
  • Say whether you need the full page or the visible viewport, if the server exposes that choice.
  • Ask the agent to close the browser even if navigation or capture fails, if the server supports cleanup in an error path.

Tool availability is not universal. Consult the selected server’s current reference for screenshot arguments, dimensions, output format, storage location, and session lifecycle controls.

5. Direct Selenium alternative: runnable Python example

If you prefer explicit control over the browser flow, this Python script opens Chrome, waits for the document to load, saves a screenshot, prints the title, and quits. Install Selenium with python -m pip install selenium, ensure Chrome is installed, and run it with python capture.py. Current Selenium includes Selenium Manager for driver discovery and caching, so a separate driver-manager package or hard-coded driver path is usually unnecessary.

from pathlib import Path
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait

url = "https://example.com"
output = Path("screenshot.png")

driver = webdriver.Chrome()
try:
    driver.set_window_size(1440, 1000)
    driver.get(url)
    WebDriverWait(driver, 20).until(
        lambda browser: browser.execute_script("return document.readyState") == "complete"
    )
    print(f"Page title: {driver.title}")
    driver.save_screenshot(str(output))
    print(f"Saved screenshot to {output.resolve()}")
finally:
    driver.quit()

The document-ready condition means the initial document has loaded; it does not guarantee that client-rendered content, fonts, or lazy images are ready. For a page-specific condition, wait for an element that you verified in the running application:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

WebDriverWait(driver, 20).until(
    EC.visibility_of_element_located((By.CSS_SELECTOR, "main"))
)

Place that wait after driver.get(url) and before save_screenshot. Replace main with a selector that exists on the target page. Selenium recommends explicit waits for the condition needed by the next action rather than relying on fixed sleeps.

6. Screenshot behavior and practical options

The exact options available through MCP depend on the server. Before building a workflow around an option, confirm that the server exposes it and check how the result is returned.

Need What to verify
Viewport or full page Whether the screenshot tool captures the visible viewport, the whole page, or both. Selenium’s basic save_screenshot captures the current browser window; full-page behavior may require additional browser-specific handling.
Size and scale Whether the tool accepts window dimensions, device scale, or a device preset. In a direct script, set the window size before capture.
Page readiness Whether the server supports waiting for a selector, navigation state, or delay. Prefer a condition tied to the page over a fixed sleep.
Output format and destination Whether it returns image content, a path, or another reference, and which formats or file paths it supports.
Browser and session lifecycle Which browser names are accepted and how to start, reuse, and stop a session. Close sessions after capture to release resources.

For dynamic pages, a successful navigation does not necessarily mean the visible state is final. Wait for the content you need, and capture only after any required interaction or overlay handling.

7. Troubleshooting

Symptom Likely cause Fix
The Selenium tools do not appear in the client The server entry is in the wrong client configuration location, the client has not reloaded, or npx failed to launch. Check the client’s current MCP configuration instructions, confirm Node.js is available to the client process, and inspect the client or server startup error.
The package fails to start Node.js is below the version required by the selected project, package resolution failed, or the repository’s setup has changed. Check the server’s current prerequisites and install instructions. The gaforov README reviewed for this guide lists Node.js 20+.
Browser or driver startup fails A supported browser is missing, browser installation is incomplete, or driver discovery cannot complete. Install a browser supported by that server, check its logs, and verify network or local environment restrictions affecting Selenium Manager. Avoid adding a separate driver manager unless the current Selenium setup requires it.
The screenshot is blank or shows the wrong page Navigation failed, the page is still rendering, or a redirect or interstitial appeared. Ask the agent to read the current URL and title, then wait for a page-specific selector before capture. Save a screenshot at the failure point to inspect what the browser displayed.
The screenshot is cut off The tool captured only the viewport or the browser window is smaller than expected. Check whether the server supports full-page capture. Otherwise set a larger window for viewport capture or use a browser-specific full-page method in a direct script.
An element click is intercepted A cookie banner, modal, or other overlay covers the target. Inspect a screenshot from the failure moment, identify the overlay, and handle it deliberately before retrying. Do not guess a selector without checking the live page.
The agent uses an unknown tool name or parameter Server tools differ, or the installed package changed. Inspect the server’s current tool list and documentation, then ask the agent to use the exposed tool and its documented arguments.
The browser remains open after an error The workflow stopped before its cleanup tool ran. Use a cleanup step in the agent instructions or a finally block in a direct Selenium script so the session is closed on failure.

8. Performance, reliability, and cost

Capture time depends on browser startup, network response, page scripts, and how long the workflow waits for the desired state. Reusing a session may avoid repeated startup work if the server supports it, but close sessions when finished. Keep waits bounded and tied to the next action so a stalled page does not hold the workflow indefinitely.

For repeatable results, use a stable URL and a verified readiness condition. A screenshot taken at the moment of failure can expose a consent banner or overlay behind a failed interaction. Selenium’s guidance recommends running an individual test multiple times before treating it as free of races. If you need a deterministic capture, account for changing page content and the target site’s access behavior.

The Selenium MCP examples use local software configuration; the research dossier does not establish a price for these community projects. Account for your own browser runtime and infrastructure if you run captures in a hosted environment. An MCP connection does not guarantee that a target site will load or permit automation.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets from over 60 known platforms are removed before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.

For more capture options and parameter details, see the ScreenshotNeo API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo supports full-page captures with lazy images loaded, selector-based element captures, dark mode, device presets and custom viewports, retina scale, PDF settings, HTML or CSS to image, custom CSS and JavaScript, click-before-capture, selector hiding, selector or network-idle waits, request and resource blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, image resizing, configurable cache TTL, signed links, asynchronous jobs with signed webhooks, batches of up to 100 URLs, a usage API, and an OpenAPI spec. Parameter names used by other screenshot APIs also work to make switching easier. All features are available on every plan.

Plans are Free for 1,000 screenshots per month with no card; Starter is $5 for 3,000; Growth is $15 for 15,000; Pro is $39 for 60,000; Scale is $99 for 250,000; and Business is $249 for 1,000,000. Yearly billing gives two months free. See ScreenshotNeo for the product details. Sign up for 1,000 free screenshots a month with no card.

10. FAQ

Is Selenium MCP an official Selenium feature?

No. Selenium’s project describes community MCP servers as third-party projects. Check the maintenance, permissions, and documentation of the specific server you install.

Can I use a different AI agent or MCP client?

Yes, if the client supports MCP and can launch or connect to the selected server. Configuration steps vary by client.

Why ask the agent to read the page title?

A title check is a quick way to confirm the browser reached a page before saving an image. For more confidence, also inspect the current URL or wait for a page-specific element.

Should I use an MCP server or a Selenium script?

Use MCP when you want an agent to invoke browser actions conversationally. Use a script when you want direct, reviewable control or need the capture inside an automated test.