ScreenshotNeo

BlogAI agents

How to Capture a Webpage Screenshot with an AI Agent

Capture a webpage with Playwright MCP or a Playwright script. Choose viewport, full-page, or element screenshots, save the image, and troubleshoot common issues.

By the ScreenshotNeo team4 October 20268 min read

To capture a webpage with an AI agent, have it open the page in a browser and call Playwright MCP’s browser_take_screenshot tool. Use the default for the visible viewport, fullPage: true for the full scrollable page, or target to capture one element. To automate captures in code instead, navigate with Playwright and call page.screenshot().

This guide is in English, including its examples. The title’s “in Hindi” describes the search topic; the steps and code work regardless of the language used to prompt an agent.

1. Capture a screenshot through an AI agent

Your agent must have access to a browser controlled through Playwright MCP, and the browser must be on the page you want to capture. Ask the agent to navigate to the URL if it is not already open, wait until the content you need is visible, and then save the screenshot.

For example, you can ask: “Open https://example.com, wait for the page content to appear, then save a full-page screenshot as webpage.png.” The MCP tool call is conceptually:

{"fullPage": true, "filename": "webpage.png"}

The documented tool is browser_take_screenshot. It captures the current page, so opening the right URL and confirming the intended state are part of the workflow. See the Playwright MCP documentation for the tool’s parameters.

Choose the capture scope

What you need What to request Important detail
What is currently visible Default screenshot Captures the viewport, not content farther down the page.
The whole scrollable page fullPage: true Cannot be combined with target.
One component or region target set to a unique CSS selector or page reference Use a selector or reference that identifies the intended element.

Example calls for each scope:

// Visible viewport
{"filename": "viewport.png"}

// Full scrollable page
{"fullPage": true, "filename": "full-page.png"}

// One element: use a unique selector or a reference from the page snapshot
{"target": "main article", "filename": "article.png"}

The target value above is an illustrative selector; replace it with a selector that matches the page you are capturing, or use the exact element reference exposed by your MCP browser session. The tool does not combine an element target with full-page capture.

Pick a filename, format, and scale

Set filename when you want a predictable output name. Playwright MCP supports PNG, JPEG, and WebP; it can infer the image type from the filename extension. If you omit the type and the extension does not determine it, the documented default is PNG. The scale option chooses CSS-pixel or device-pixel output: CSS-pixel scale is the compact option, while device-pixel scale produces more device pixels and can create a larger image. Use the scale values supported by your installed MCP server version.

{"filename": "page.webp", "type": "webp"}

If you only need to inspect the result, a tool response may make the screenshot visible to the agent. If you need a reusable artifact, specify a filename and check the path or output directory used by your MCP setup; relative paths are resolved by the server, so they may not be relative to your shell’s current directory.

2. Capture with Playwright code

Use Playwright’s Page API when you want the capture in a repeatable script, need to control browser setup, or want to integrate the image into a build or review workflow. Install Playwright and its browser for your environment following the official installation guide.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  try {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    await page.goto('https://example.com', { waitUntil: 'load' });
    await page.screenshot({ path: 'webpage.png', fullPage: true });
  } finally {
    await browser.close();
  }
})();

Save this as screenshot.js and run it with Node.js after installing Playwright and its browser. The example uses a fixed viewport for repeatability and closes the browser even if navigation or capture fails. Playwright documents page.screenshot() options, including the path and fullPage options, in its Page API reference.

Wait for the content you need

A navigation event does not guarantee that every image, client-rendered section, or delayed component is ready. If the page has a known landmark, wait for it explicitly before capturing:

await page.goto('https://example.com', { waitUntil: 'load' });
await page.locator('main').waitFor({ state: 'visible' });
await page.screenshot({ path: 'webpage.png', fullPage: true });

Replace main with a selector that identifies the content you need. For a viewport capture, omit fullPage or set it to false. For an element capture, use the locator’s screenshot method:

await page.locator('article').screenshot({ path: 'article.png' });

Selectors must match the page’s actual structure. If the selector matches several elements, narrow it to the intended one. When an agent is choosing controls or discovering page structure, use its accessibility snapshot or page inspection tool alongside the screenshot; the screenshot itself is an image, not a structured interaction map.

3. When to use an agent, MCP, or a script

Approach Best fit What to keep in mind
AI agent with Playwright MCP Ad hoc capture based on a natural-language request, or a page already open in the agent’s browser The agent needs a configured MCP browser and access to the correct page and output location.
Playwright script Repeatable captures, custom browser setup, or integration into an existing Node.js workflow You manage the browser installation, script, navigation, waiting, and output path.
Screenshot API A simple HTTP request from an application, script, or automation without managing a browser session Check the API’s options, response behavior, and billing rules for the service you choose.

Playwright MCP’s screenshot tool offers a practical agent-driven route. Its documented choices cover the viewport, a target element, or the full scrollable page. For a workflow that instead needs a screenshot API, ScreenshotNeo is an option described below.

4. Or skip the browser setup

If you need a screenshot from code or an agent workflow without setting up a browser, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF. The examples below use the API base URL from its documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Replace YOUR_API_KEY with your key and change the target URL. The Node.js example expects a runtime with built-in fetch and uses top-level await in an ES module.

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

5. Troubleshooting

Symptom Likely cause What to do
The screenshot shows the wrong page The agent captured its current tab or session state. Ask it to navigate to the exact URL, verify the page, then capture.
Content is missing or still loading The capture happened before a delayed or client-rendered section appeared. Wait for a visible landmark or the specific element, then take the screenshot.
The page is cut off The capture used the viewport default. Set fullPage: true in MCP or fullPage: true in page.screenshot().
The element capture fails or selects the wrong region The selector is absent, ambiguous, or the target reference is stale. Inspect the current page, use a unique selector or current page reference, and retry. Do not combine target with fullPage.
The output file is missing The filename is relative to the MCP server’s output or working directory, or the script wrote to another directory. Check the MCP server configuration and process working directory; use an explicit path where supported.
The file extension and image encoding do not match The requested format and filename disagree. Use a matching extension and type, or let the MCP tool infer type from the extension.
The capture is unexpectedly large Full-page dimensions or device-pixel scale increased output size. Capture only the viewport or target when sufficient, and use CSS-pixel scale for a compact image.
The page is blank or shows a challenge The site may require a session, block automation, or fail to render in the current setup. Confirm the URL and browser session, wait for content, and follow the site’s access rules. Do not assume a screenshot tool can bypass a challenge.

6. Performance, reliability, and cost

  • Capture only what you need. Viewport or element screenshots can avoid the work and file size of a tall full-page image. Full-page capture is useful when content below the fold matters.
  • Wait on a meaningful condition. A specific visible element gives the script or agent a clear readiness signal. A fixed delay can be simpler but may waste time or still be too short when the page is slow.
  • Keep captures reproducible. Use a consistent viewport, URL, browser context, and filename pattern when comparing captures over time. Authenticated pages require an appropriate session in the browser context.
  • Handle failures at the caller. Scripts should surface navigation and file errors; API clients should use timeouts and check HTTP status before treating a response as an image.
  • Account for where work runs. MCP output paths belong to the server environment, while script paths belong to the script process. In remote or containerized setups, retrieve files from the environment that created them.
  • Check the pricing model before scaling. Playwright self-hosting uses your own browser runtime and compute. ScreenshotNeo offers 1,000 free shots per month without a card, then paid plans from $5 for 3,000; its billing response headers identify whether a shot was billed.

7. FAQ

Can an AI agent screenshot a page that is already open?

Yes, if its browser MCP session controls that page. Ask it to capture the current page and specify the desired scope and output filename.

Does a screenshot let the agent click or understand page controls?

No. It gives visual information. Use the browser’s accessibility snapshot or interaction tools when the agent needs structure or actionable element references.

Can I capture just one element and the full page in the same MCP call?

No. The documented Playwright MCP parameters do not allow target and fullPage together. Capture the element or the full page separately.

Which format should I choose?

PNG is the documented fallback; JPEG and WebP are also supported by the MCP screenshot tool. Choose a matching filename extension, and consider whether your downstream tool accepts that format.

Does full-page mean a PDF?

No. Full-page capture produces an image of the scrollable page. Use a PDF capture workflow when the deliverable must be a PDF.

References