How to Capture Website Screenshots with an AI Agent and Save Them to S3
Capture a website with an AI agent using Playwright, upload the image bytes to S3, and choose the right approach for page screenshots, desktop capture, or session replay.
Direct answer: use browser automation to navigate to the intended page and capture screenshot bytes, then upload those bytes explicitly to an Amazon S3 object. Saving an image locally and storing it in S3 are separate operations. For a page viewport, element, or full-page image, Playwright over Amazon Bedrock AgentCore Browser is a practical route. Use AgentCore’s OS-level InvokeBrowser screenshot action when you need a full desktop view or content outside the browser viewport. If you want replayable interaction history rather than just an image, configure AgentCore session recording separately.
This guide uses Python, Playwright, and boto3 for the complete upload path. It assumes you have an AgentCore Browser tool available and AWS credentials configured for your runtime. AgentCore’s documented Playwright integration connects to a managed browser over Chrome DevTools Protocol (CDP). AWS: Using AgentCore Browser with Playwright
1. Choose what the screenshot should contain
| Capture type | Use it for | How to capture |
|---|---|---|
| Viewport | The visible browser page at a particular size and scroll position | page.screenshot() |
| Element | A chart, card, or other specific component | page.locator("selector").screenshot() |
| Full page | The full scrollable webpage | page.screenshot(full_page=True) |
| Full desktop | The browser plus visible desktop content outside its viewport | AgentCore InvokeBrowser screenshot action |
Viewport, element, and full-page capture are browser-level operations. Playwright documents these capture modes and can return the image as bytes, which can be passed directly to another service. Playwright: Screenshots AgentCore’s CDP-based browser connection is suited to webpage navigation and DOM interaction; InvokeBrowser provides OS-level actions for cases such as native dialogs, alerts, or full-desktop screenshots. AWS: Browser OS action
2. Set up the browser and AWS access
- Create or choose an AgentCore Browser tool in the AWS account and region where you will run the capture.
- Configure AWS credentials for the process using an IAM role or another standard AWS credential provider. The identity that uploads the screenshot needs
s3:PutObjectaccess to the destination object or prefix. - Install Python dependencies in a virtual environment.
python3 -m venv .venv
source .venv/bin/activate
pip install bedrock-agentcore playwright boto3
On Windows, activate the environment with .venv\\Scripts\\activate. This script uses the built-in AgentCore browser identifier aws.browser.v1; if you are using a custom browser, set BROWSER_ID to its browser identifier. AWS documents session creation, viewport settings, timeouts, and session shutdown in its browser session guide. AWS: Managing Browser Sessions
3. Capture and upload with Python
Save this as capture_to_s3.py. The screenshot is kept in memory as bytes and uploaded directly. Set the required environment variables before running it.
import asyncio
import os
import re
from urllib.parse import urlparse
import boto3
from bedrock_agentcore.tools.browser_client import browser_session
from playwright.async_api import async_playwright
REGION = os.environ.get("AWS_REGION", "us-west-2")
BROWSER_ID = os.environ.get("BROWSER_ID", "aws.browser.v1")
BUCKET = os.environ["S3_BUCKET"]
URL = os.environ["TARGET_URL"]
MODE = os.environ.get("CAPTURE_MODE", "full_page") # viewport, full_page, or element
SELECTOR = os.environ.get("CAPTURE_SELECTOR")
def object_key(url: str) -> str:
host = urlparse(url).netloc.replace(":", "_") or "site"
slug = re.sub(r"[^A-Za-z0-9._-]+", "-", urlparse(url).path.strip("/"))[:80]
return f"website-shots/{host}/{slug or 'homepage'}.png"
async def main():
s3 = boto3.client("s3", region_name=REGION)
# The context starts and stops the managed browser session.
with browser_session(REGION, identifier=BROWSER_ID) as client:
ws_url, headers = client.generate_ws_headers()
async with async_playwright() as p:
browser = await p.chromium.connect_over_cdp(ws_url, headers=headers)
try:
if not browser.contexts:
raise RuntimeError("AgentCore browser returned no browser context")
context = browser.contexts[0]
page = context.pages[0] if context.pages else await context.new_page()
response = await page.goto(URL, wait_until="domcontentloaded", timeout=60000)
if response and response.status >= 400:
raise RuntimeError(f"Navigation returned HTTP {response.status}: {URL}")
# Wait for the page's visible content to settle. Prefer a real
# page-specific selector when the site exposes one.
await page.locator("body").wait_for(state="visible", timeout=15000)
await page.wait_for_timeout(500)
if MODE == "viewport":
image_bytes = await page.screenshot(type="png")
elif MODE == "full_page":
image_bytes = await page.screenshot(type="png", full_page=True)
elif MODE == "element":
if not SELECTOR:
raise ValueError("Set CAPTURE_SELECTOR when CAPTURE_MODE=element")
target = page.locator(SELECTOR).first
await target.wait_for(state="visible", timeout=15000)
image_bytes = await target.screenshot(type="png")
else:
raise ValueError("CAPTURE_MODE must be viewport, full_page, or element")
key = object_key(URL)
s3.put_object(
Bucket=BUCKET,
Key=key,
Body=image_bytes,
ContentType="image/png",
)
print(f"Uploaded s3://{BUCKET}/{key} ({len(image_bytes)} bytes)")
finally:
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Run it after setting the bucket and URL:
export AWS_REGION=us-west-2
export S3_BUCKET=my-screenshot-bucket
export TARGET_URL=https://example.com
python capture_to_s3.py
For a component capture, set CAPTURE_MODE=element and CAPTURE_SELECTOR to a CSS selector. For just the current viewport, set CAPTURE_MODE=viewport. The default is a full-page PNG. The script derives a simple key from the host and URL path; for production, prefer an application-owned identifier or add a timestamp or unique job ID to avoid overwriting earlier captures.
4. Make the capture match the page state
A screenshot is only useful if it represents the intended page state. Navigation completion does not guarantee that client-rendered content, fonts, images, or data have finished loading. Use the narrowest wait condition that matches the page:
- Wait for a specific component:
await page.locator(".report-chart").wait_for(state="visible"). - Wait for a known application state:
await page.wait_for_function("() => window.reportReady === true"). - Wait for network quiet only when appropriate:
await page.goto(URL, wait_until="networkidle"). Pages with analytics, polling, or streaming may never become network-idle. - Wait for lazy content: full-page capture does not guarantee every site lazy-loads images correctly. Scroll in increments to trigger loading, then wait for images or target content before taking the final capture.
Keep browser navigation and screenshot work deterministic: use a fixed viewport, explicit locale/time zone if your setup exposes those options, and a stable test account or public page. Avoid waiting on arbitrary long delays when a selector or application-ready signal exists. For pages whose content changes continuously, decide whether the goal is a snapshot at a chosen instant or a settled page, then define that condition explicitly.
5. Uploading and delivering the S3 object
put_object stores the selected image bytes at the bucket and key you specify. Set an accurate content type such as image/png, image/jpeg, or image/webp. Keep the bucket private unless public access is an explicit requirement. Limit IAM permissions to the intended bucket and prefix, and do not put AWS credentials in page scripts, prompts, browser-visible storage, or URLs.
If a frontend needs temporary access to the image, return a presigned URL from your application instead of making the bucket public. The URL is a bearer link: anyone who receives it can use it until it expires, so use an appropriately short expiration and avoid logging or publishing it. AWS’s architecture example uses a presigned URL to let a client fetch screenshot data from S3. AWS: AgentCore Browser quickstart
6. When to use AgentCore session recording
Session recording serves a different purpose from uploading one screenshot. When enabled on a custom AgentCore browser, it stores a broader interaction trail in S3, including DOM changes, user actions, console messages, CDP events, and network events. Use it when debugging or auditing a browser session; use the explicit screenshot upload flow when your application needs a chosen image object.
Recording requires an S3 bucket, a browser execution role with the required S3 write permissions, and a custom browser configured with recording enabled. Review the recorded data for sensitive page content and network information before enabling it. AWS: Session Recording and Replay
7. Full-desktop screenshot with InvokeBrowser
Playwright captures page content. For screenshots that must include the complete browser desktop or content outside its viewport, use AgentCore’s OS-level screenshot action. The API request is SigV4-signed; AWS documents awscurl usage and requires the session ID header. The following is the request shape; substitute your region and active session ID.
awscurl -X POST \
"https://bedrock-agentcore.us-west-2.amazonaws.com/browsers/aws.browser.v1/sessions/invoke" \
-H "Content-Type: application/json" \
-H "Accept: application/json" \
-H "x-amzn-browser-session-id: YOUR_SESSION_ID" \
--service bedrock-agentcore \
--region us-west-2 \
-d '{
"action": {
"screenshot": { "format": "PNG" }
}
}'
Inspect the response format for the SDK/API version in use, decode its PNG payload to bytes, then upload those bytes with the same S3 put_object operation shown above. InvokeBrowser returns a screenshot result; it does not by itself place that image in your application’s S3 key. Only one action is supplied per request. AWS: Browser OS action
8. Reliability, performance, and cost
- Keep capture scope reasonable. Full-page images can be much larger and slower to encode and transfer than viewport or element images. Capture only what the consumer needs.
- Use idempotent keys where retries matter. A stable job ID makes a retried upload replace the same object; unique IDs preserve each attempt. Decide which behavior your workflow needs.
- Retry transient failures selectively. Retry timeouts, throttling, or transient service errors with bounded backoff. Do not repeatedly retry permanent navigation errors, invalid selectors, access-denied responses, or a target that intentionally blocks automation.
- Always close the browser connection and session. Use cleanup in
finallypaths. AWS advises stopping sessions when done to release resources and avoid unnecessary charges. - Watch image and storage growth. Choose an S3 lifecycle/retention policy that matches how long screenshots are useful, especially when capturing frequently or retaining full-page images.
- Budget the whole pipeline. Browser session/runtime usage, S3 storage, requests, and data transfer are separate considerations. Consult current AWS pricing and quotas for your region and configuration; no fixed performance or cost estimate applies to every page and session.
9. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Playwright cannot connect or reports WebSocket errors | Wrong region, missing AWS credentials/permissions, expired session, or mismatched browser identifier | Check the region and browser ID, verify AWS credentials, ensure the session is active, and follow the current AgentCore Playwright setup instructions. |
AccessDenied during upload |
The caller lacks s3:PutObject permission on the key, or a bucket policy/KMS policy denies it |
Grant narrowly scoped access to the target bucket and prefix; check encryption-key permissions if using a customer-managed key. |
| Object exists but the browser cannot display it | The bucket is private, the caller has no read permission, or the object metadata is wrong | Use an authorized application path or presigned URL; set the correct ContentType. |
| Screenshot is blank or content is missing | Capture happened before rendering, an element was not visible, the page failed, or the wrong capture scope was chosen | Check navigation status, wait for a page-specific ready selector, verify the target, and compare viewport with full-page capture. |
| Element locator times out | Selector changed, element is in a frame, or the page state was not reached | Inspect the page structure, wait for the right state, and use the correct frame/locator for embedded content. |
| Full-page image omits lazy images | The site loads images only after they approach the viewport | Scroll through the page to trigger lazy loading, then wait for the required images before capturing. |
| Recording is absent from S3 | Recording is not enabled on a custom browser, the execution role lacks permissions, or bucket/prefix configuration is wrong | Check the browser configuration, role trust and S3 permissions, bucket/prefix, and AgentCore logs. AWS troubleshooting guide |
InvokeBrowser returns validation or access errors |
Invalid action/session, out-of-range coordinates for coordinate actions, inactive session, or missing permission | Check the active session and request shape, verify IAM permission for bedrock-agentcore:InvokeBrowser, and use a screenshot to confirm the viewport before coordinate actions. |
10. Or skip the browser setup
If you only need a website image file, ScreenshotNeo is a website screenshot API and MCP server. Make the call from your application, then upload the returned bytes to S3 with your normal AWS SDK or CLI. The API returns an image or PDF response; this code writes the response body as an image file, which you can then upload to S3.
See the ScreenshotNeo API documentation for request options. Example with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. The API response can be uploaded to S3 as bytes using your application’s S3 client.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does an agent framework have to call the browser?
No. The AgentCore Browser can be controlled with Playwright directly; an LLM or agent framework is optional when the capture steps are known in advance.
Does S3 store the screenshot automatically?
No. A local save or browser screenshot produces image data. Your code must upload it, or you must configure a separate service such as AgentCore session recording for its own recording artifacts.
Can I use JPEG or WebP instead of PNG?
Yes, where supported by the capture path. Match the S3 ContentType to the encoded format and use a file extension that matches it.
Can the same workflow save a PDF?
Browser screenshot calls return images. PDF output is a separate browser operation and should be uploaded with application/pdf; use the relevant browser/API documentation for the exact options in your chosen runtime.


