How to Capture a Website Screenshot with an AI Agent Using a Proxy
Configure Playwright to capture a website through a proxy, give an AI agent bounded browser tools, and troubleshoot screenshot and proxy failures.
To capture a website screenshot with an AI agent through a proxy, configure a browser automation tool such as Playwright with the proxy endpoint, then let the agent call a small set of browser actions: navigate to an allowed URL, wait for the page state you need, and capture the viewport, an element, or the full page. The proxy routes browser requests; it does not grant access to a site or authorize bypassing its access controls.
The agent, browser, proxy, and screenshot operation are separate parts. The agent plans and calls tools, Playwright performs deterministic browser operations, and the proxy routes requests according to its configuration. The example below is an illustrative baseline based on documented Playwright APIs; it was not run as part of the research. Adapt and verify it with your Playwright version, browser, operating system, and proxy provider.
1. Install Playwright and prepare the proxy
Use a proxy endpoint supplied by your proxy operator. Confirm its scheme (for example, HTTP or SOCKS), hostname, port, and whether it requires a username and password. Install Playwright and its browser:
python -m pip install playwright
python -m playwright install chromium
Playwright documents proxy configuration on browser launches and contexts, including a server, optional bypass domains, and optional credentials. See the BrowserType API reference and its browser installation guidance.
Keep proxy credentials in environment variables or a secret store. Do not put them in agent instructions, source control, logs, or screenshot artifacts. The scheme and authentication format are provider-specific; a proxy URL that works for one provider may not work for another.
2. Capture a page with Python and Playwright
This script demonstrates the browser layer. It reads proxy settings from the environment, opens a page, waits for the document load state, and saves a screenshot. It takes one URL argument and does not implement an AI model or agent loop: expose the navigation and capture operations as bounded tools for your agent to call.
import asyncio
import os
import sys
from urllib.parse import urlparse
from playwright.async_api import async_playwright
async def main():
if len(sys.argv) != 2:
raise SystemExit("Usage: python capture.py https://example.com")
target_url = sys.argv[1]
parsed = urlparse(target_url)
if parsed.scheme not in ("http", "https") or not parsed.hostname:
raise SystemExit("Target must be an http:// or https:// URL")
proxy_server = os.environ.get("PROXY_SERVER")
proxy = None
if proxy_server:
proxy = {"server": proxy_server}
username = os.environ.get("PROXY_USERNAME")
password = os.environ.get("PROXY_PASSWORD")
if username:
proxy["username"] = username
if password:
proxy["password"] = password
bypass = os.environ.get("PROXY_BYPASS")
if bypass:
proxy["bypass"] = bypass
async with async_playwright() as p:
browser = await p.chromium.launch(
headless=True,
proxy=proxy,
)
try:
page = await browser.new_page(
viewport={"width": 1440, "height": 900},
device_scale_factor=1,
)
response = await page.goto(
target_url,
wait_until="domcontentloaded",
timeout=45000,
)
# Wait for a meaningful state if the page has one. Replace this
# with a site-specific selector when possible.
await page.locator("body").wait_for(state="visible", timeout=15000)
await page.screenshot(path="screenshot.png", full_page=True)
status = response.status if response else "no main response"
print(f"Captured {page.url} (HTTP status: {status}) to screenshot.png")
finally:
await browser.close()
asyncio.run(main())
Set the endpoint and optional credentials through your shell, then run the script:
export PROXY_SERVER='http://proxy.example:8080'
export PROXY_USERNAME='your-username'
export PROXY_PASSWORD='your-password'
python capture.py https://example.com
For a SOCKS proxy, use the scheme and endpoint format documented by your provider. Playwright’s API reference includes HTTP and SOCKS proxy examples. Avoid embedding secrets in shell history on shared systems; use your platform’s secret management instead.
3. Give the agent a bounded browser tool
Do not ask an agent to invent arbitrary browser actions when the task only needs a screenshot. Give it narrow tools such as navigate(url), wait_for_selector(selector), and capture_screenshot(scope, selector). Validate the URL against an explicit allowlist before navigation, set timeouts, and return the final URL, response status when available, and artifact path along with the image.
- Agent: interprets the request, chooses the permitted URL and capture scope, and decides whether a visual check is needed.
- Playwright: navigates, waits for a page state, performs permitted interactions, and captures the image.
- Proxy: routes requests as configured. It does not ensure that the destination will load or that access is permitted.
- Artifact handler: stores or returns the image and reports which URL and page state were captured.
In production instructions, identify the target URL and permitted domains, the desired output path or delivery method, and what the agent should report if navigation or capture fails. Keep proxy credentials outside agent-visible inputs and outputs. A screenshot is for visual inspection; use page locators or snapshots to identify elements for interaction. The Playwright screenshots documentation describes viewport, element, and full-page captures and distinguishes screenshots from interaction references.
4. Choose viewport, element, or full-page capture
| Scope | Use it for | Playwright option |
|---|---|---|
| Viewport | The currently visible screen, useful for a visual check at a fixed window size. | page.screenshot(path="shot.png") |
| Element | A particular card, chart, or other selected region. | page.locator("#target").screenshot(path="element.png") |
| Full page | The whole scrollable document, useful for a page overview. | page.screenshot(path="full.png", full_page=True) |
Full-page capture and selected-element capture are different scopes; do not combine them in one screenshot request. Large documents can produce tall, memory-heavy images. For a full-page result, consider whether the page needs to be scrolled first to load lazy images, and whether a long image is practical for your agent or downstream system.
For consistent results, set the viewport and device scale factor explicitly. Responsive layout, fonts, animation, personalized content, and dynamic widgets can change the image. Wait for a meaningful selector or application-ready condition instead of relying on an arbitrary delay. If you need a cookie banner, login state, or post-interaction state included, specify that in the task and perform the relevant action before capturing.
5. cURL, Python, and Node.js alternatives
If the goal is simply to obtain a screenshot over an HTTP API, a browser agent and proxy setup may be unnecessary. ScreenshotNeo is a website screenshot API and MCP server: one GET request can return a PNG, JPEG, WebP, or PDF. Its API can accept custom headers and other capture options; see the ScreenshotNeo API documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
These examples use an API key placeholder; keep the real key in secret storage and do not expose it in browser output or client-side code. For public image embedding, consult the docs for signed links. ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
6. Or skip the browser setup
Use the ScreenshotNeo API when you need a screenshot without operating a browser process and proxy yourself. One GET request returns an image or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the API documentation for request options. Cookie banners are accepted like a visitor and removed before the shot, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
7. Hosted browser options and when to use them
Local Playwright gives your application control over the browser process and local output files. Hosted execution can make sense when the browser session should run remotely. OpenAI’s computer-use documentation describes an OpenAI-hosted browser environment controlled through application events and screenshot observations when available. The Browser Use project documentation describes a hosted sandbox and proxy options; check its current project guidance for details.
These hosted options do not necessarily have the same session controls, data handling, proxy behavior, geographic coverage, or pricing. Before choosing one, check who configures and operates the proxy, supported protocols and locations, credential handling, screenshot delivery and retention, rate limits, cost, and operational control. A configured proxy is not a guarantee that a target site will load.
8. Reliability, performance, and cost
- Wait for state, not time alone: use a page-specific selector or readiness condition where possible. A fixed sleep may be too short on a slow page and wasteful on a fast one.
- Bound each operation: set navigation and selector timeouts, close the browser in a
finallyblock, and return a clear failure result to the agent. - Reduce image size when appropriate: choose the smallest viewport and device scale that meets the visual task. Full-page images and high-density captures consume more memory and produce larger artifacts.
- Control nondeterminism: use a consistent viewport, browser version, locale, and page state. Dynamic pages, animations, ads, and personalized content may vary between runs.
- Plan for proxy failures: a proxy can add connection latency or fail independently. Record safe diagnostic details such as the target hostname and error category, but redact proxy credentials and sensitive headers.
- Account for total cost: local runs use your own compute and proxy service; hosted browser and proxy pricing depends on the provider and plan. Review limits and current terms directly rather than assuming a particular cost or performance.
- For repeatable QA: Playwright Test can save screenshots on failure and optionally record video and traces. These artifacts are configurable and documented as off by default in its test configuration guide.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Proxy connection refused or timed out | Wrong endpoint, port, protocol, network route, or unavailable proxy. | Confirm the scheme and host with the provider, verify network access from the browser host, and test a short navigation timeout before retrying. |
| Proxy authentication error | Missing or incorrect credentials, or the provider expects a different authentication setup. | Check the provider’s required fields and pass credentials via environment or secret storage. Do not print them in logs. |
| Browser launches but the page does not load | The proxy may be misconfigured, the destination may be unavailable, or the site may reject the request. | Inspect the navigation exception and response status; verify the endpoint and target independently. Do not treat proxy routing as a way to evade access restrictions. |
| TLS or certificate error | A corporate network may intercept TLS, or the target certificate may be invalid. | Use the organization’s approved trusted CA configuration and follow browser guidance for that environment. Do not disable certificate checks as a general fix. |
| Screenshot is blank or incomplete | Capture happened before content rendered, navigation failed, or content is loaded only after scrolling or interaction. | Check the final URL and response status, wait for an application-specific selector, and perform only the required scroll or interaction before capture. |
| Images are missing from a full-page capture | Lazy-loaded assets may not have been requested because their regions were never brought into view. | Scroll through the page to trigger loading, wait for the relevant assets, then capture. Verify the resulting page state rather than assuming every site loads images the same way. |
| Element screenshot times out | The selector does not match, is hidden, or appears only after a delayed state change. | Inspect the selector, wait for it to become visible, and capture the element alone. Fall back to a viewport capture if the requested target is absent. |
| Flaky visual results | Fonts, animations, responsive layout, or dynamic content differ between runs. | Pin viewport and browser settings, wait on meaningful content, and use Playwright Test traces or failure screenshots to inspect the failing run. |
| Agent reports the wrong page as captured | Redirects or navigation errors changed the final location. | Return page.url and response status with the artifact; validate the final host against the allowed domain policy. |
10. FAQ
Can an AI agent take a website screenshot?
Yes. The agent can decide when to request a capture, while a browser automation layer such as Playwright performs the navigation and screenshot operation.
Does using a proxy guarantee the site will load?
No. The proxy controls routing; availability, authentication, site policy, and the target’s own responses still determine whether navigation succeeds.
Can I capture the full page and a selected element at once?
No. Choose one scope for each capture request. Take separate screenshots if both artifacts are required.
Should I use local Playwright or a hosted browser?
Use local Playwright when you want to control the browser process and local files. Consider hosted execution when remote browser operation fits your deployment, after checking the provider’s proxy, session, artifact, and cost details.


