How to Use Gemini CLI and MCP to Screenshot a Website
Use Gemini CLI with Chrome DevTools MCP to capture a website, inspect the returned image, and save a screenshot file. Includes setup, browser-agent options, and troubleshooting.
To screenshot a website with Gemini CLI, connect it to a browser-control MCP server, ask Gemini to open the target URL and take a screenshot, then verify whether the tool returned an image, saved a file, or did both. The two documented routes are configuring Chrome DevTools MCP yourself and enabling Gemini CLI’s browser agent, which manages its bundled Chrome DevTools MCP server.
What you need
- Gemini CLI installed and working.
- A Chrome installation compatible with the selected browser-control route.
- The target website URL and, if you want a file, a destination path that exists or can be created.
- Permission to run the MCP tool actions requested by Gemini CLI. Review each approval prompt before allowing an action.
Browser versions, authentication, package behavior, and container support can change. Check the current Gemini CLI MCP documentation, browser-agent documentation, and instructions for the MCP server you choose before relying on a particular setup.
Route A: configure Chrome DevTools MCP directly
This route gives you an explicit MCP server configuration. The Google Codelab demonstrates a server named chrome-devtools that runs chrome-devtools-mcp@latest, with a Chrome executable path and headless mode. The exact config file location and schema can depend on your Gemini CLI version, operating system, and chosen server version; use the current Gemini CLI and server documentation for the config format.
1. Install and locate Chrome
Install Chrome and identify its executable path on your machine. Do not copy a path from another operating system without checking that it exists locally. The Codelab’s example uses a configured Chrome executable and the --headless option; whether that option is suitable depends on your environment and server version.
2. Add the MCP server configuration
In Gemini CLI’s MCP configuration, add a server entry named chrome-devtools that runs the server package and passes the options required by your installed Chrome. The Codelab illustrates the shape of this setup, but its path and flags are examples, not universal values. Follow the server’s current instructions for package execution, Chrome path, headless support, and configuration syntax.
3. Start Gemini CLI and check the connection
gemini
At the Gemini CLI prompt, list the MCP servers and tools:
/mcp list
Confirm the Chrome DevTools server is available and that its browser tools appear. If it does not appear, see Troubleshooting.
4. Ask for a screenshot and a file path
Replace the example URL and path with your own. If the site is a local app, start it in a separate terminal first and use its actual local address.
Open https://example.com with the configured Chrome DevTools MCP tools. Wait for the page to finish loading, take a screenshot, and save it as output/website.png. Tell me whether the screenshot was returned as an image, where the file was saved, and whether saving succeeded.
When Gemini asks to approve a browser action, inspect the requested action and approve it only if it matches your task. Then check that output/website.png exists and open it to confirm it contains the expected page. A natural-language request does not guarantee that a file was written.
Route B: use Gemini CLI’s browser agent
The browser-agent route uses a bundled chrome-devtools-mcp server that Gemini CLI launches automatically when enabled. This can be convenient when you want the browser agent’s managed interaction workflow instead of manually configuring a server. Enable browser_agent in Gemini CLI’s settings.json using the current documented settings format, then start Gemini CLI and ask it to use the browser agent to capture the target page.
The browser-agent documentation describes three session modes:
| Mode | What it means | Use it when |
|---|---|---|
persistent |
The default mode; the browser profile is retained. | The task needs browser state retained between sessions. |
isolated |
A temporary browser profile is used and discarded after the session. | You want a temporary, separate profile for a task. |
existing |
The agent attaches to an existing Chrome instance with remote debugging enabled. | You need to work with an already-running Chrome session. |
Check the live browser-agent documentation for exact settings, Chrome requirements, and attachment steps. At the time of the cited documentation, Chrome 144 or later is required. Optional visual-agent support can use visualModel for screenshot-based visual identification; its authentication requirements and availability differ by setup, and the documentation says it is unavailable with “Sign in with Google.” Treat these as version-sensitive details and verify them before configuring a workflow.
Prompt for a saved screenshot
Use the browser agent to open https://example.com. Wait until the page is ready, take a screenshot, save it to output/website.png, and confirm the full path and whether the file was created.
For a local application, replace the URL with the address printed by your development server. If you need an authenticated view, choose a session mode and profile that has the required login state; avoid placing passwords or secrets in the prompt.
Image returned by MCP versus screenshot saved to disk
MCP tool results can contain text and image blocks. An image block may include a type, base64-encoded image data, and a MIME type such as image/png. Gemini CLI can make the image available to the model and display it separately. That response image is not, by itself, evidence that a file was written to your filesystem.
| What you need | What to request | How to verify |
|---|---|---|
| Gemini to inspect the screenshot | Ask the browser tool to capture the page and inspect or describe the result. | Check that the tool result contains image content and Gemini can discuss the visible page. |
| A screenshot file | Specify an exact output path and ask Gemini to confirm the write. | Check that the path exists and open the image with a viewer. |
| Both | Ask for the capture, file save, and confirmation of both outputs. | Verify the returned image and the saved file separately. |
Choosing a route
| Consideration | Direct Chrome DevTools MCP | Browser agent |
|---|---|---|
| Setup | You configure the MCP server and browser executable. | Enable the agent in settings; Gemini CLI launches its bundled server. |
| Control | Useful when you want to manage the server entry and its options directly. | Useful when you want the managed browser-agent workflow. |
| Browser state | Depends on the server configuration and browser profile you use. | Offers persistent, isolated, and existing session modes. |
| Environment constraints | Check package, browser path, headless mode, and remote environment support. | Check current Chrome minimum version, settings, and authentication requirements. |
Options for a more reliable capture
- Wait for the right condition: Ask the tool to wait for the page or a specific visible element to load. A page’s initial navigation completing does not always mean client-rendered content is ready.
- Use a stable viewport and state: For repeatable captures, use the same browser mode, viewport, login state, and target URL each time, where the selected tool exposes those controls.
- Use a local URL that the browser can reach: A browser running inside a container or remote environment may not share your host machine’s
localhost. Use a reachable address for that environment. - Choose a file path explicitly: Use a relative path such as
output/website.pngonly when you know the CLI’s working directory. Ask for the resolved path and confirm it exists. - Be deliberate with session modes: Persistent profiles retain browser state; isolated profiles are temporary; existing mode relies on a Chrome instance configured for remote debugging. Use the mode that fits the site’s authentication and your privacy needs.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
/mcp list does not show the server |
Invalid config, package startup failure, or a missing executable. | Check Gemini CLI’s current MCP configuration format, the server command, package availability, and Chrome path. Restart the CLI after changing configuration. |
| Chrome does not launch | The executable path is wrong, Chrome is absent, or the selected headless/remote option is unsupported in the environment. | Confirm the browser is installed and the path is valid. Compare the flags with the server’s current instructions and try in a supported local environment. |
| The page is blank or incomplete | The app is still rendering, a resource failed, or the browser cannot reach the URL. | Open the URL in the same browser environment, wait for a meaningful page element, and verify the app server is running and reachable there. |
| The screenshot appears in the response but there is no file | The tool returned MCP image content without writing a local file. | Ask explicitly for a file at a precise path, then check that path in the CLI’s working directory. Use a tool action that supports saving if needed. |
| File is saved somewhere unexpected | A relative path was resolved from a different working directory. | Ask for the absolute resolved path or use a known writable destination; then inspect the file directly. |
| Login state is missing | The selected browser profile is new or isolated. | Use the appropriate retained or existing session mode, or authenticate in the browser without putting credentials in prompts. Follow your organization’s security policy. |
| Browser agent reports an unsupported version | The installed Chrome is below the currently documented minimum. | Check the live browser-agent docs and update Chrome if required by that version of the agent. |
| Visual identification does not work | Optional visual-agent configuration or authentication is missing or unsupported. | Check the current visualModel and authentication requirements. The documented Google sign-in limitation may apply. |
| Works locally but not in Cloud Shell or a container | The remote environment may not support the browser process, GUI/headless behavior, or network route the workflow needs. | Review the Codelab and server environment requirements. The cited Codelab says BrowserMCP does not work in Google Cloud Shell for its exercise; use a supported environment or a reachable browser host. |
Performance, reliability, and cost
Capture time depends on the website, network, browser startup, and how long the page takes to render. Browser-based capture also consumes local or remote browser resources. For repeatable work, avoid arbitrary short waits when the page has a clear readiness signal, and verify output files rather than assuming a successful tool response means a successful save.
Gemini CLI, the browser, and the MCP server are separate components, so failures can come from configuration, browser startup, site loading, tool approval, or file writing. Check each layer in that order. This workflow uses the browser and agent setup you provide; the cited sources do not establish a per-screenshot price for Gemini CLI or Chrome, so account for any separate infrastructure or model costs applicable to your setup.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. For an API capture, one GET request returns an image or PDF; its MCP server provides screenshot tools for AI agents. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
FAQ
Can Gemini CLI screenshot a localhost app?
Yes, if the browser launched by your selected MCP route can reach that address. In container or remote setups, the browser’s localhost may refer to the container or remote host rather than your development machine.
Does the MCP image automatically become a PNG file?
No. Image content returned to Gemini and a screenshot written to disk are distinct outputs. Request a destination path and verify the file.
Can I use a logged-in website?
Use a browser session that has the needed state, such as an appropriate retained profile or an existing Chrome session. Follow the current browser-agent instructions and avoid sharing credentials in prompts.
Where can I find the exact Gemini CLI configuration syntax?
Use the official MCP server guide and the current CLI commands reference; configuration details can change between versions.


