How to Use an AI Agent to Capture Webpage Screenshots in WebP Format
Use Playwright CLI, MCP, or JavaScript to capture a webpage as WebP. Choose the right page state, capture area, scale, and quality.
To capture a webpage screenshot as a real WebP image with an AI agent, use a screenshot tool that supports WebP and set its format explicitly. With Playwright, you can use the agent CLI, the Playwright MCP tool, or the Page API. For the Page API, call page.screenshot({ path: "page.webp", type: "webp" }). A .webp filename can also help select the format in interfaces that infer it from the extension, but explicitly setting the type avoids ambiguity.
Decide first whether you need the visible viewport, a particular element, or the entire scrollable page. Then wait for the page state you need, set WebP, choose an image scale and quality, and have the agent report where it saved the result. Playwright defaults can be PNG when neither type nor extension selects another format, so do not assume a screenshot is WebP just because an agent took it.
1. Choose an agent workflow
Use the interface that fits how your agent works: a command-line session, an MCP connection, or code that calls the Playwright Page API.
| Workflow | Best fit | Format and output | Capture scope |
|---|---|---|---|
| Playwright agent CLI | An agent that can run shell commands in a browser session | Set --type=webp and --filename |
Viewport or full page; the CLI also supports a target |
| Playwright MCP | An MCP-connected agent that invokes browser tools | Pass type: "webp"; the screenshot tool can return the image inline |
Viewport, target, or full page; target and full page cannot be combined |
| Playwright Page API | An agent running JavaScript with a Playwright page | Set type: "webp" and a .webp path |
Viewport, full page, or an element screenshot |
These interfaces expose different controls and output handling. Match examples to the Playwright version and interface you have installed, since documentation and tool behavior can change.
2. Capture a WebP with the Playwright agent CLI
Open the target page in the CLI browser session, wait for the content you intend to capture, then run:
playwright-cli screenshot --type=webp --full-page --filename=page.webp
This command requests a full-page WebP file named page.webp. To capture the current viewport, remove --full-page. The CLI supports --type=png, --type=jpeg, or --type=webp, plus --filename and --hires. Use the target option documented for your CLI version when you need a particular page element.
If you rely on extension inference instead of --type, name the output with a .webp extension. Explicit type selection is clearer, particularly in scripts and agent instructions.
3. Capture WebP through Playwright MCP
Ask the connected agent to call browser_take_screenshot with WebP selected. For example, an MCP client can pass arguments in this shape:
{
"type": "webp",
"fullPage": true,
"scale": "css"
}
Use fullPage: true for the scrollable page. For one element, pass the tool’s target option instead; the documentation says a target cannot be combined with full-page capture. Use scale: "device" when you want device-pixel resolution. When filename is omitted, the tool can return the image inline in its response as well as saving it according to its documented behavior. If you need a file for a later job, provide a filename where the tool supports it and have the agent report the resulting path.
4. Capture WebP with the Playwright Page API
In JavaScript code where page is an existing Playwright Page, this captures the full page and writes a lossless WebP:
await page.goto("https://example.com");
await page.screenshot({
path: "page.webp",
type: "webp",
quality: 100,
fullPage: true,
});
For a viewport screenshot, omit fullPage. For an element, call the locator’s screenshot method, for example:
await page.locator("main article").screenshot({
path: "article.webp",
type: "webp",
quality: 90,
});
Choose a selector that identifies the intended element. If it matches multiple elements, is hidden, or is not yet rendered, the capture may fail or target the wrong content; wait for the intended locator and make the selector specific.
5. Set page state, capture area, scale, and quality
Wait for the content you actually need
There is no universal wait duration that guarantees a useful screenshot. A static page may be ready after navigation; a client-rendered page, dashboard, or image-heavy article may need a specific element or content state to appear. Wait for a meaningful selector or condition where possible. A fixed delay can help with known animation or delayed content, but it may waste time on fast pages and still be too short on slow ones.
Before capture, consider whether the screenshot should show consent dialogs, transient loading states, animations, or overlays. If your task is to document the page as a visitor sees it, those may be relevant. If you need the underlying page content, decide how your workflow should handle them instead of assuming every page has the same state.
Choose viewport, element, or full page
- Viewport: captures the currently visible area. Use it for a fold-specific view or a compact preview.
- Element: captures a selected element, such as an article, chart, or product card. Confirm it is visible and uniquely identified.
- Full page: captures content beyond the viewport. Long pages can produce very tall images and take more memory to process.
Playwright CLI, MCP, and Page API use different controls for these scopes. In MCP, full-page capture and a target are mutually exclusive according to the screenshot tool documentation.
Choose scale
CSS scale produces one image pixel per CSS pixel. Device scale uses device pixels, which can create larger dimensions on high-DPI displays. Use CSS scale for predictable dimensions tied to the page layout; use device scale if you need the higher-resolution rendering and can handle the additional pixels. The CLI calls its high-resolution option --hires; the MCP screenshot tool exposes scale.
Choose WebP quality
For the Playwright Page API, WebP quality 100 is lossless; lower values are lossy. Use lossless when preserving rendered detail is more important than file size. Lower quality can reduce output size, but the result depends on the page and settings, so measure your own captures before choosing a value. Do not assume a particular percentage reduction.
6. Make the capture repeatable
- Open the intended URL in the agent’s browser session.
- Wait for the content or page condition that defines a ready capture.
- Choose viewport, element, or full-page scope.
- Set WebP explicitly and specify an output filename or path if the interface supports it.
- Set scale and, for the Page API, quality according to fidelity and file-size needs.
- Have the agent report the filename and path, then inspect the image when correctness matters.
A useful agent instruction states the URL, what content must be visible, the required capture scope, the WebP output path, and whether the agent should return the image or save it for downstream processing. This makes the result easier to review and avoids an agent silently returning a PNG or capturing the wrong page state.
7. cURL, Python, and Node.js for ScreenshotNeo
If you want an API call instead of managing a browser session, ScreenshotNeo accepts one GET request with a URL and returns a screenshot or PDF. For the WebP workflow, use the format option documented in the ScreenshotNeo API documentation alongside the URL and access key. These examples use the supplied base request; add the documented WebP parameter for the API version you use.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
These base examples show request and file handling. Consult the documentation for the precise format parameter, supported response behavior, and available options before relying on an extension alone. Check the response headers and saved file when your pipeline requires a particular format.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its API takes a URL in one GET request; its MCP tools let Claude, Cursor, and other MCP clients take screenshots, get page information, or capture PDFs.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for format and capture options. Before the shot, ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| The output is PNG instead of WebP | The tool defaulted to PNG because neither type nor extension selected WebP. | Set type: "webp" or --type=webp, and use a .webp filename. |
| The screenshot is blank or shows a loading state | The page had not rendered the desired content when capture ran. | Wait for a meaningful selector or page condition, then capture again. |
| The full-page image is unexpectedly large | The page is long, device scale is enabled, or both. | Use viewport or element capture, choose CSS scale, or reduce WebP quality if lossy output is acceptable. |
| The wrong element is captured or capture fails | The selector is ambiguous, hidden, or not present yet. | Use a more specific locator and wait for it to become visible. |
| MCP rejects the screenshot arguments | Full-page capture was combined with a target, or the argument names differ from the installed tool version. | Use either fullPage or target, and check the MCP screenshot docs for the connected version. |
| The agent cannot find the saved file | The output path is relative to a different working directory or the tool returned the image inline. | Use an explicit path where supported and ask the agent to report it. For MCP, check the tool’s documented save behavior. |
Performance, reliability, and cost
Full-page screenshots and device-scale captures contain more pixels than viewport captures, so they can take more memory and produce larger files. Element captures can keep output focused. Lower WebP quality may reduce size, while quality 100 in the Page API is lossless. Measure representative pages and inspect image quality at the sizes your downstream system will use; the cited documentation does not establish a universal speed or file-size advantage.
Reliability depends on reaching the intended page state before capture. Use a selector or condition tied to required content, make the output path explicit, and inspect important results. Playwright’s cited docs describe format and capture controls but do not establish a cross-browser guarantee or a universal output-file validation command. For API cost, ScreenshotNeo lists 1,000 free shots monthly with no card, then Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Clean shots are billed, while bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing.
FAQ
Does a .webp filename guarantee WebP output?
Format inference depends on the interface. Explicitly choose WebP with the documented type option when available.
Can an agent return the screenshot instead of saving it?
Yes. Playwright MCP can return the image inline when filename is omitted; use a saved path when another process needs a file.
Is WebP quality 100 lossy?
For Playwright’s Page API, quality 100 is documented as lossless. Lower values are lossy.
Can I capture a full page and a particular element at once?
The Playwright MCP screenshot tool does not combine a target with full-page capture. Choose the scope that matches the task.


