How to Use an AI Agent to Screenshot a Website with a Custom User Agent
Set a custom user agent before navigation, choose the right screenshot target, and understand what the override can—and cannot—change.
To screenshot a website with an AI agent and a custom user agent, configure the browser context with the desired user-agent string before creating or navigating the page. Then navigate, wait for the page state you need, and capture the viewport, a selected element, or the full page. The user-agent override changes the browser identity string; it does not guarantee a particular layout or bypass bot protection.
This guide uses Playwright because its browser context exposes a userAgent setting and its screenshot APIs support viewport and full-page capture. The same sequence applies when an agent’s browser integration exposes equivalent context settings. If the integration does not let you configure the browser before navigation, this method may not be available in that integration.
1. Decide what browser identity and layout you need
Use a custom user-agent string when you specifically need a particular string sent by the browser. If you are trying to reproduce a device layout, also set the viewport or choose an appropriate device preset. A user agent, viewport, and broader device emulation are related but distinct settings. Playwright device presets can include viewport and other emulation values; its Desktop Chrome preset uses a Windows-specific user agent. The [Playwright emulation guide](https://playwright.dev/docs/emulation) documents these options.
Use a real, intended user-agent value from the system or test scenario you are reproducing. A made-up or mismatched string can produce confusing results, and changing the string alone does not make the browser behave like a complete physical device.
2. Configure Playwright before navigation
Install Playwright and its Chromium browser in your project using the [official Playwright installation instructions](https://playwright.dev/docs/intro). Save the following as screenshot.mjs, replace the example user-agent with the one required by your test, and run it with Node.js in an environment where Chromium can launch.
import { chromium } from 'playwright';
const url = process.argv[2] ?? 'https://example.com';
const userAgent = process.env.CUSTOM_USER_AGENT;
if (!userAgent) {
throw new Error('Set CUSTOM_USER_AGENT to the user-agent string you want to use.');
}
const browser = await chromium.launch();
try {
const context = await browser.newContext({
userAgent,
viewport: { width: 1280, height: 800 }
});
const page = await context.newPage();
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.screenshot({ path: 'page.png', fullPage: true });
await context.close();
} finally {
await browser.close();
}
Run it by setting the environment variable to your chosen string:
CUSTOM_USER_AGENT='Mozilla/5.0 (compatible; ScreenshotAgent/1.0)' node screenshot.mjs https://example.com
The example string is illustrative, not a recommendation for impersonating a browser. For repeatable work, keep the intended value in configuration or a secret store rather than hard-coding it in source control.
Wait for the state the screenshot needs
The example waits for domcontentloaded, which means the initial document has been parsed; it does not mean that every image, API request, or client-rendered component is ready. Choose the wait condition based on the page and the content you need:
- For content available in the initial document, a document navigation milestone may be enough.
- For a specific client-rendered component, wait for that component’s selector to become visible.
- For a known animation or delayed banner, use a deliberate short delay only when needed.
- Avoid assuming that network idle is appropriate for every site; pages with polling or persistent connections may never become idle.
For example, replace navigation and capture with a wait for a known element:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.locator('[data-testid="report-ready"]').waitFor({ state: 'visible', timeout: 15000 });
await page.screenshot({ path: 'report.png' });
Use selectors that belong to the page you control or are documented for your test. A selector that does not exist will time out; handle that as a page-state failure rather than silently capturing an incomplete result.
3. Choose the screenshot target
Playwright supports screenshots of the current viewport, a specific element, and the full scrollable page. Its [screenshot documentation](https://playwright.dev/mcp/tools/screenshots) also describes screenshot options for agent workflows.
Viewport screenshot
Capture only what is visible in the current viewport. This is useful for a fixed screen state or a visual comparison at a known width and height.
await page.screenshot({ path: 'viewport.png' });
Element screenshot
Capture one component, such as a chart or card, using a locator. Make sure the element is attached and visible before capture.
const chart = page.locator('#revenue-chart');
await chart.waitFor({ state: 'visible' });
await chart.screenshot({ path: 'chart.png' });
Full-page screenshot
Capture the page beyond the current viewport:
await page.screenshot({ path: 'full-page.png', fullPage: true });
Full-page capture can be large and may expose layout problems on pages that rely on sticky elements or lazy loading. If below-the-fold images have not loaded, scroll through the page or wait for the relevant content before capturing, then inspect the result.
Image format and device scale
Playwright’s screenshot API supports image output options such as PNG, JPEG, and WebP where documented by the installed API version. Set the type explicitly when your downstream system expects a particular format. Device scale affects pixel dimensions; high-resolution device pixels do not use the same coordinate scale as CSS pixels for agent mouse actions. See the [Playwright Screenshots and PDF command reference](https://playwright.dev/agent-cli/commands/screenshots-pdf) for the agent CLI’s screenshot options and coordinate warning.
await page.screenshot({ path: 'page.webp', type: 'webp', fullPage: false });
Do not assume that setting a high device scale changes CSS layout. It changes raster output resolution; use viewport and device emulation settings for layout behavior.
4. Use an agent CLI or MCP browser tool
If your agent controls Playwright through its CLI or MCP integration, configure the context through that integration before navigating. The exact mechanism depends on the runtime. The Playwright agent CLI documents commands including screenshot, screenshot [target], screenshot --filename=<f>, screenshot --type=<png|jpeg|webp>, and screenshot --full-page. These capture commands do not themselves establish that a custom user agent was configured; set the context first using the integration’s supported configuration path.
For agent tasks that need to understand page structure or read content, pair the image with a DOM-oriented inspection or accessibility snapshot. Playwright describes screenshots as useful for visual layout, canvas or chart content, and bug documentation; accessibility snapshots are more useful for structure, text, and interaction references. A screenshot alone is not a reliable substitute for structured page information.
5. Other implementation paths
Puppeteer: Puppeteer documents Page.screenshot() for capture. Its page API can be used in a browser workflow that sets the desired user agent before navigation. This dossier does not establish a performance or image-fidelity winner between Puppeteer and Playwright; choose based on the integration and browser setup your agent already supports. See [Puppeteer’s screenshot API](https://pptr.dev/api/puppeteer.page.screenshot).
Hosted browser: A hosted browser can avoid managing a local browser installation and may suit remote or repeated capture workloads. Cloudflare Browser Run documents screenshots, browser sessions, Playwright integration, and AI-agent browsing. Its documentation specifically cautions that a custom user agent does not bypass its bot identification: Browser Run requests are identified as bot traffic even when a custom value is set. Review the current [Browser Run documentation](https://developers.cloudflare.com/browser-run/) and [pricing page](https://developers.cloudflare.com/browser-run/pricing/) before choosing it; service limits and prices can change.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The page still shows a desktop layout | The user-agent string changed, but the viewport or device settings did not. | Set a viewport appropriate to the target and, when needed, use a documented device preset. Confirm which settings the preset supplies. |
| The server still identifies the request as automated | A user-agent override does not bypass bot protection. Hosted browser providers may explicitly identify their traffic as bot traffic. | Do not treat the override as a way around site controls. Use an authorized test environment or the site’s supported access method. |
| The screenshot is blank or missing a component | The page may render content asynchronously, a navigation may have failed, or the capture happened before the component appeared. | Check navigation errors and page state; wait for the specific expected selector, then capture again. Use an accessibility snapshot or DOM inspection to distinguish missing content from a visual-only issue. |
| A wait times out | The selector is wrong, the target state never occurs, or the page is still loading indefinitely. | Verify the selector and expected state, inspect the page, and use a timeout suited to the operation. Avoid unbounded waits and indiscriminate network-idle waits. |
| Full-page output misses lazy images | Those images load only after their section enters the viewport. | Scroll through the page or otherwise trigger the relevant content, wait for it to load, and then capture. |
| Agent click coordinates do not line up with the image | The screenshot uses device pixels while interaction coordinates use CSS pixels. | Account for the device scale factor or use element locators for interactions where possible. |
| Chromium will not launch | The runtime may lack the installed browser or required system dependencies. | Install the browser required by the Playwright setup for that environment, or use a hosted browser integration supported by the agent. |
7. Performance, reliability, and cost
For a local Playwright run, the main practical costs are browser setup, browser process resources, navigation time, and the bytes written or passed back to the agent. Reuse a browser process for batches of captures when your runner supports it, while creating a separate context for each distinct user-agent or isolated session. Keep screenshots to the smallest target and dimensions that satisfy the task; full-page, high-resolution images require more output data and can take longer to produce or transmit.
Reliability depends on the target page and chosen readiness condition. Use finite navigation and selector timeouts, record failures distinctly from successful images, and avoid treating a timed-out or blocked page as a valid screenshot. A user-agent override is a rendering input, not a guarantee of server response, access, or content negotiation.
Local Playwright has no hosted browser usage meter, though it uses your own compute and maintenance time. Hosted browser services have provider-specific limits and charges. Cloudflare’s current pricing documentation lists browser-time allowances and additional charges; check that page at the time you deploy because terms can change. No neutral benchmark in the reviewed sources establishes one local or hosted option as universally faster or cheaper.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns a screenshot or PDF, and its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents. The API accepts parameter names used by other screenshot APIs, which can make switching easier. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Replace YOUR_API_KEY and the target URL with your values. The custom-user-agent workflow above is for browser contexts you control; use ScreenshotNeo’s documented parameters for the capture options it supports.
- Cookie banners are accepted and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses indicate the page verdict and billing status in headers.
- An MCP server lets AI agents take screenshots, get page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
9. FAQ
Does a custom user agent make an agent browser undetectable?
No. It changes the configured user-agent string. It does not guarantee human classification or override bot controls.
Should I use a device preset or set only the user agent?
Set only the user agent when the string itself is what you need to reproduce. Choose a device preset or configure additional emulation values when you need a broader device-like setup, and verify the viewport and other values it applies.
Can the agent use a screenshot to extract page text?
It can inspect visible text in the image, but screenshots do not provide the same structured content and interaction references as a DOM or accessibility snapshot. Use the representation that fits the task.
Can I use this workflow with a hosted Playwright browser?
Yes, if that service and agent integration let you set the context user agent before navigation. Check the provider’s own documentation for its limits and bot-identification behavior.


