How to Store Scheduled Website Screenshots in Amazon S3
Schedule browser captures with EventBridge and Lambda, upload private screenshots to S3, and manage retries, monitoring, retention, and costs.
Use Amazon EventBridge Scheduler to invoke an AWS Lambda function on a recurring schedule. The function renders the page in a browser, uploads the screenshot bytes to a private Amazon S3 bucket, and logs whether rendering and upload succeeded. Add retry and failure handling, an S3 Lifecycle rule for retention, and an alert based on capture success or staleness.
The schedule invocation alone does not prove a screenshot was saved. Verify both the page capture and the S3 write. Browser automation packaging and compatibility depend on the Lambda runtime and deployment setup; choose and validate that runtime for your workload.
1. Choose the capture and storage design
Before writing code, decide what each capture represents and how it should be retrieved.
- Capture scope: a viewport screenshot for a stable visual checkpoint, or a full-page screenshot when content below the fold matters. Specify the viewport and page readiness conditions so captures are comparable.
- Cadence: create one schedule per cadence, or group sites that share a cadence and compatible execution requirements.
- Object key: use a timestamped key for a history of captures, such as
site-id/2026/10/04/15-30-00Z.png. Use UTC to avoid daylight-saving ambiguity. A fixed key is useful for a latest-image pointer but overwrites the previous object unless versioning is enabled. - Event configuration: pass a site identifier and URL, or a configuration key. Keep credentials and sensitive configuration out of the event payload when possible; store secrets separately and grant the function scoped access.
- Access: keep the bucket private and give readers authenticated, scoped access. A screenshot may expose account details, unpublished material, or personal information.
AWS recommends EventBridge Scheduler for scheduled targets. It supports recurring cron and rate schedules and invokes Lambda asynchronously. The Scheduler execution role authorizes invocation; the Lambda execution role authorizes the function’s access to AWS resources. Keep these roles separate and scoped to their respective tasks. See AWS guidance for invoking Lambda on a schedule and the Lambda execution role documentation.
2. Create the S3 bucket and permissions
Create a bucket in the region appropriate for your workload, block public access, and decide whether versioning is needed. Configure the Lambda function’s execution role with only the required S3 upload permissions, restricted to the bucket and object prefix it writes. Include the logging permissions needed for its CloudWatch logs. The Scheduler execution role needs permission to invoke the Lambda target; it does not need the function’s S3 permissions.
S3 applies server-side encryption with S3 managed keys (SSE-S3) to new uploads by default. AWS describes this as the base level of encryption for every bucket. Encryption does not replace access control: keep the bucket private and scope reader access. Use a customer-managed KMS key if you need its additional control and policy requirements; check current key charges and permissions before choosing it. See S3 default encryption documentation.
3. Implement the capture function
The following Python handler shows the event contract, timestamped object key, browser capture boundary, S3 upload, and structured outcome logging. It uses Playwright as the browser automation interface and boto3 for S3. The AWS references for this guide do not establish which Playwright package, browser binary, Lambda layer, or container image works for a particular runtime. Package a compatible browser runtime and dependencies for your selected Lambda environment, then verify its memory and time limits before deployment. The handler is the application code; it is not a complete browser deployment package.
Example event payload:
{
"site_id": "docs-home",
"url": "https://example.com",
"full_page": true,
"viewport": {"width": 1440, "height": 1000}
}
Set the Lambda environment variable SCREENSHOT_BUCKET to the private bucket name. Install boto3 and a Playwright version with a browser binary compatible with the selected Lambda runtime in the deployment artifact or image.
import json
import logging
import os
from datetime import datetime, timezone
from urllib.parse import urlparse
import boto3
from playwright.sync_api import sync_playwright
logger = logging.getLogger()
logger.setLevel(logging.INFO)
s3 = boto3.client("s3")
BUCKET = os.environ["SCREENSHOT_BUCKET"]
def lambda_handler(event, context):
site_id = event["site_id"]
url = event["url"]
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError("url must be an absolute http or https URL")
viewport = event.get("viewport", {"width": 1440, "height": 1000})
full_page = bool(event.get("full_page", True))
captured_at = datetime.now(timezone.utc)
timestamp = captured_at.strftime("%Y/%m/%d/%H-%M-%SZ")
key = f"{site_id}/{timestamp}.png"
try:
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport=viewport)
response = page.goto(url, wait_until="networkidle", timeout=45000)
if response is not None and response.status >= 400:
raise RuntimeError(f"Page returned HTTP {response.status}")
image = page.screenshot(full_page=full_page, type="png")
browser.close()
s3.put_object(
Bucket=BUCKET,
Key=key,
Body=image,
ContentType="image/png",
)
logger.info(json.dumps({
"site_id": site_id,
"url": url,
"captured_at": captured_at.isoformat(),
"render_result": "success",
"s3_result": "success",
"bucket": BUCKET,
"key": key,
}))
return {"statusCode": 200, "bucket": BUCKET, "key": key}
except Exception as exc:
logger.exception(json.dumps({
"site_id": site_id,
"url": url,
"captured_at": captured_at.isoformat(),
"render_result": "failed",
"s3_result": "not_confirmed",
"error": str(exc),
}))
raise
Adjust the readiness strategy to the site. networkidle can be unsuitable for pages with persistent connections or continuous polling; a bounded delay or waiting for a meaningful selector may work better. If the site requires authentication, provide credentials through a protected mechanism and make sure the resulting screenshot’s storage and reader permissions are appropriate. Do not put secrets in the S3 key.
4. Configure EventBridge Scheduler
Create a recurring schedule with a cron expression for a calendar-based cadence or a rate expression for a fixed interval. Set the target to the capture Lambda function and pass a JSON event with the site identifier and URL (or a configuration reference). Configure the schedule’s execution role to invoke that function.
For production, configure the Scheduler retry policy and failed-event retention to match the outage and recovery window you can tolerate. Use a dead-letter queue if failed invocations should be inspected or replayed. Scheduler invokes Lambda asynchronously, so the function’s own outcome must be monitored separately from the invocation.
Keep schedule configuration manageable: one schedule per site is straightforward when cadences differ; a shared schedule can pass a list or configuration key when sites share a cadence and can be processed within the function’s execution limits. If a batch contains multiple sites, record each site’s outcome independently so one failure does not obscure which captures succeeded.
5. Validate a capture before relying on the schedule
- Invoke the Lambda directly with a test event and confirm the function can launch its packaged browser.
- Check the CloudWatch log entry for the site identifier, render result, upload result, and S3 key.
- Confirm the object exists in the private bucket and has the expected content type and image dimensions.
- Invoke the schedule or wait for its next run, then verify both the invocation and the resulting object.
- Test a deliberately invalid URL or inaccessible target in a nonproduction configuration to confirm failures are logged and handled as expected.
AWS documents checking CloudWatch logs to confirm Scheduler invoked the Lambda function. That confirms invocation, not successful rendering or upload. Instrument and alert on the application-level outcomes as well. See the scheduled invocation documentation.
6. Set retention and deletion behavior
Use an S3 Lifecycle rule to expire screenshots after the period they remain useful, or transition older captures to another storage class after evaluating that class’s costs and access behavior. Lifecycle rules apply to existing as well as future objects. Transitions and some storage classes have request costs or minimum-duration considerations, so evaluate the current rules for your region and access pattern. See S3 Lifecycle management.
If bucket versioning is enabled, expiration of the current object can add a delete marker while earlier noncurrent versions continue to consume storage. Add a noncurrent-version expiration action if the intention is to remove old screenshot bytes after the retention period. Lifecycle actions are asynchronous, so deletion may not happen at the exact instant the configured age is reached. Review version and delete-marker lifecycle behavior in AWS Lifecycle troubleshooting.
| Need | Key and bucket choice | Retention consideration |
|---|---|---|
| Historical visual record | Unique UTC timestamp in each key | Expire old objects and, if versioned, noncurrent versions |
| Addressable latest screenshot | Fixed key, such as site-id/latest.png |
Overwrites replace the current object; versioning can preserve prior versions and add storage cost |
| Recovery from accidental deletion | Versioned bucket with restricted access | Set noncurrent-version expiration deliberately; current-object expiration alone may not remove old bytes |
7. Reliability, performance, and cost
Reliability and missed captures
- Set retry limits and failed-event retention for the schedule based on the recovery window. Use a dead-letter queue when operators need to inspect or replay failures.
- Emit structured logs for each site with the requested URL, capture timestamp, render status, upload status, and object key.
- Alert on failed captures and on a stale last-success timestamp. Invocation metrics alone cannot detect a page that rendered blank or an S3 write that failed.
- Use timestamped keys for history. If a fixed key is required, decide whether S3 Versioning is needed to recover overwritten images.
Performance
Browser startup and page rendering are generally the work that determines execution time for this architecture. Select Lambda memory and timeout settings using representative pages and the chosen browser package; this dossier does not provide benchmark figures or a verified runtime configuration. Bound navigation and readiness waits, and avoid waiting for network idle on pages that never become idle. For batches, account for the slowest page and keep the workload within the function’s timeout and memory limits.
Cost
There is no reliable dollar estimate here: total cost depends on region, capture cadence, browser execution time and resources, Lambda invocations, S3 requests, storage duration, transitions, retrieval, and version retention. Estimate those components using current AWS pricing for your region. Reduce unnecessary captures, set a useful retention period, and ensure Lifecycle rules also clean up noncurrent versions when appropriate. A transition is not automatically cheaper for short-lived or frequently retrieved screenshots.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Scheduler ran but no object appeared | The log confirms invocation only; rendering or upload may have failed. | Inspect function logs for render and S3 outcomes, then check the function role’s bucket and prefix permissions. |
| Browser executable not found or launch fails | The browser binary or native dependencies are absent or incompatible with the Lambda runtime/package. | Package a browser runtime compatible with the selected runtime and architecture. Validate it in the deployed environment before relying on scheduled runs. |
| Function times out during page load | Navigation or readiness condition takes too long, or the page maintains network activity. | Use a bounded timeout and an appropriate readiness condition; increase the function timeout only within the workload’s operational constraints. |
| Screenshot is blank or incomplete | The capture happened before useful content rendered, or the site returned an interstitial/error page. | Wait for a meaningful selector or a suitable bounded delay, inspect the page response, and log a capture outcome that distinguishes a rendered page from a successful upload. |
| AccessDenied on S3 upload | The Lambda execution role lacks the required write permission or is scoped to a different bucket/prefix; encryption policy requirements may also apply. | Grant only the necessary object write action to the intended prefix and review bucket and KMS policies if using a customer-managed key. |
| Old screenshots remain after expiration | Lifecycle processing is asynchronous, or versioning retains noncurrent versions. | Allow for Lifecycle processing delay and add noncurrent-version expiration when permanent removal is intended. |
| One site’s failure hides other results | A shared batch reports only one overall result. | Log and track status per site, and structure batch handling so each site’s outcome is observable. |
| Large history costs more than expected | Captures or noncurrent versions are retained longer than intended, or transitions/retrieval add charges. | Review cadence, object age, versioning, lifecycle actions, and current regional pricing together. |
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Make a GET request with a URL to receive a PNG, JPEG, WebP, or PDF, then store the response bytes in S3 from your scheduled function. The call below captures a page; your scheduled worker still needs to upload the returned bytes to S3.
See the ScreenshotNeo API documentation for request options. Example cURL call:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For a scheduled S3 workflow, your function can make the same request, check the response and billing headers, and pass the response bytes to s3.put_object. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card.
FAQ
Should every website have its own schedule?
Only when sites need different cadences or execution settings. Compatible sites can share a schedule and be passed as a batch or configuration reference, provided each site’s result is tracked separately.
Does a successful Lambda invocation mean the screenshot is safe to use?
No. Confirm that the page rendered as intended and that the object upload succeeded. A function can be invoked successfully while returning a blank or unsuitable page.
Can an S3 event start another workflow after upload?
Yes. S3 can send supported events to EventBridge for downstream processing such as thumbnail generation. This is optional for scheduled capture; see S3 EventBridge notifications.
What should be checked before choosing a browser package?
Check compatibility with the chosen Lambda runtime, architecture, browser binary, native dependencies, deployment size, memory, and timeout. Validate the packaged function in its deployed environment because the AWS references cited here do not specify a working package/version combination.


