How to Store Website Screenshot History in Amazon S3
Build a durable screenshot archive in Amazon S3 with timestamped keys, searchable manifests, versioning, lifecycle retention, and private access.
Store each website screenshot as a separate Amazon S3 object with a key that includes the site identity and capture time. Save searchable capture details—such as the original URL, UTC timestamp, viewport, format, and capture-tool version—in a JSON manifest. Keep the bucket private, enable Versioning if you need recovery from accidental changes or deletions, and use Lifecycle rules to apply a deliberate retention period.
S3 identifies an object by its bucket and key, with an optional version ID; AWS does not prescribe a screenshot-history schema. The naming pattern below is an application design choice. [Amazon S3 object model]
1. Choose a history layout
A useful layout groups records by a stable site identifier and UTC date, then gives every capture a unique timestamped name. For example:
screenshots/
docs-example/
2026/10/04/
20261004T143012Z__desktop__a1b2c3.webp
20261004T143012Z__desktop__a1b2c3.json
Use a normalized site identifier such as docs-example, not a raw hostname as an unvalidated path component. The suffix can be a random ID or capture-job ID to prevent collisions when captures happen within the same timestamp resolution. Do not put credentials, session tokens, or sensitive query parameters in an object key.
| Choice | Recommendation | Reason |
|---|---|---|
| New key per capture or overwrite one key | New key per capture | Each record is independently addressable and easy to list. Overwriting can make history depend on version-aware reads. |
| Timestamp in UTC or local time | UTC, ISO-like sortable format | Sort order is stable across regions and daylight-saving changes. |
| Metadata or manifest | JSON sidecar manifest | Supports richer fields and future schema changes than object metadata alone. |
| Versioning on or off | Enable when recovery matters | Versioning preserves earlier revisions of the same key, but it does not replace the per-capture index. |
Keep fields needed to find an image in the key or a queryable catalog. S3 prefixes support organization and listing, but S3 does not provide a general SQL search over arbitrary JSON bodies. For large archives, maintain a database or catalog indexed by site, capture time, and tags, and store the S3 key in that record.
2. Create and secure the bucket
- Create a bucket in the AWS Region appropriate for your application, residency needs, and users.
- Keep S3 Block Public Access enabled. Grant access to the application through an IAM role or narrowly scoped policy rather than making screenshot objects public. AWS recommends retaining Block Public Access. [Block Public Access]
- Use the default server-side encryption for a baseline. New S3 uploads are encrypted at rest by default with SSE-S3. Use SSE-KMS when your key policy or other KMS controls are required. Require HTTPS/TLS for requests in transit. [S3 encryption] [S3 security best practices]
- Enable Versioning if you need to restore accidental overwrites or deletions, and plan lifecycle cleanup for noncurrent versions.
- Give the capture process only the permissions it needs, such as writing under its archive prefix. Give readers separate, read-only access as appropriate.
When Versioning is enabled, overwrites create another version. A delete normally creates a delete marker, which hides the key from ordinary reads while prior versions remain. Versioning can therefore increase retained storage unless you manage noncurrent versions. [S3 Versioning]
3. Capture, upload, and write a manifest with Python
The example below captures a page with Playwright, uploads the image and a JSON sidecar using Boto3, and writes unique keys. It assumes Python 3.10 or later, an AWS identity configured through the normal SDK credential chain, and a bucket you already created. The role needs permission to put objects under the selected prefix.
python -m pip install playwright boto3
python -m playwright install chromium
import hashlib
import json
import os
import re
import uuid
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlsplit
import boto3
from playwright.sync_api import sync_playwright
BUCKET = os.environ["SCREENSHOT_BUCKET"]
URL = os.environ.get("PAGE_URL", "https://example.com")
SITE_ID = os.environ.get("SITE_ID", "example-com")
WIDTH = int(os.environ.get("VIEWPORT_WIDTH", "1440"))
HEIGHT = int(os.environ.get("VIEWPORT_HEIGHT", "900"))
# Keep identifiers predictable and safe for use in an object key.
site_id = re.sub(r"[^a-zA-Z0-9._-]+", "-", SITE_ID).strip("-.") or "site"
captured_at = datetime.now(timezone.utc)
stamp = captured_at.strftime("%Y%m%dT%H%M%SZ")
run_id = uuid.uuid4().hex[:12]
key_base = f"screenshots/{site_id}/{captured_at:%Y/%m/%d}/{stamp}__{WIDTH}x{HEIGHT}__{run_id}"
with sync_playwright() as playwright:
browser = playwright.chromium.launch()
page = browser.new_page(viewport={"width": WIDTH, "height": HEIGHT}, device_scale_factor=1)
response = page.goto(URL, wait_until="networkidle", timeout=60000)
# A response can be absent for non-HTTP schemes or some navigation cases.
status = response.status if response else None
page.screenshot(path="capture.png", full_page=True)
final_url = page.url
title = page.title()
browser.close()
image_path = Path("capture.png")
image_bytes = image_path.read_bytes()
sha256 = hashlib.sha256(image_bytes).hexdigest()
manifest = {
"schema_version": 1,
"site_id": site_id,
"source_url": URL,
"final_url": final_url,
"captured_at": captured_at.isoformat().replace("+00:00", "Z"),
"viewport": {"width": WIDTH, "height": HEIGHT, "device_scale_factor": 1},
"format": "png",
"http_status": status,
"page_title": title,
"capture_software": "Playwright Chromium",
"sha256": sha256,
"image_key": f"{key_base}.png",
}
s3 = boto3.client("s3")
s3.put_object(
Bucket=BUCKET,
Key=manifest["image_key"],
Body=image_bytes,
ContentType="image/png",
Metadata={"site-id": site_id, "captured-at": manifest["captured_at"], "sha256": sha256},
)
s3.put_object(
Bucket=BUCKET,
Key=f"{key_base}.json",
Body=json.dumps(manifest, ensure_ascii=False, separators=(",", ":")).encode("utf-8"),
ContentType="application/json",
)
print(json.dumps({"image_key": manifest["image_key"], "manifest_key": f"{key_base}.json"}))
Run it with environment values set, for example:
export SCREENSHOT_BUCKET=my-private-screenshot-archive
export PAGE_URL=https://example.com
export SITE_ID=example-com
python capture_to_s3.py
The sample waits for network idle, which is convenient for many static pages but can time out on pages with continuous requests. For those pages, use a less strict navigation condition such as domcontentloaded, then wait for a specific selector or a short, justified delay before capturing. Treat a navigation timeout as a capture failure and avoid uploading an empty or misleading image.
4. Upload an existing screenshot with AWS CLI
If another system already creates the screenshot, upload it with a unique key and attach basic searchable metadata. This command assumes the CLI has credentials and a default Region configured:
aws s3 cp ./capture.png \
s3://my-private-screenshot-archive/screenshots/example-com/2026/10/04/20261004T143012Z__desktop__a1b2c3.png \
--content-type image/png \
--metadata site-id=example-com,captured-at=2026-10-04T14:30:12Z
Upload the manifest as a second object with ContentType=application/json. Avoid relying on user metadata for a large or evolving schema; use it for a few compact lookup hints, and put the complete record in the sidecar or your catalog.
5. List and retrieve history
For a small archive, list a site prefix and sort keys by their UTC timestamp. S3 returns keys in lexicographic order for a prefix listing, so a fixed-width timestamp makes the key order useful. Large listings are paginated; use an SDK paginator rather than assuming one response contains every capture.
aws s3api list-objects-v2 \
--bucket my-private-screenshot-archive \
--prefix screenshots/example-com/2026/10/04/ \
--query 'Contents[].{Key:Key,Size:Size,LastModified:LastModified}'
To read a specific object without making the bucket public, download it with an authorized identity:
aws s3 cp \
s3://my-private-screenshot-archive/screenshots/example-com/2026/10/04/20261004T143012Z__desktop__a1b2c3.png \
./capture.png
For a private web review flow, your application can authorize a user and then issue a short-lived presigned URL. Avoid placing long-lived AWS credentials in a browser or embedding them in screenshot URLs.
6. Set retention and control archive growth
Choose retention based on the purpose of the archive and any applicable policy. There is no universal retention period for website screenshots. Use an S3 Lifecycle rule to transition older objects or expire objects when they are no longer needed. Check current storage-class behavior, retrieval requirements, and prices before choosing a transition; the research does not establish project-specific savings. [S3 Lifecycle]
For a versioned bucket, configure noncurrent-version expiration as well as current-object retention. Review expired delete marker cleanup where it applies. Otherwise, an apparently deleted key can still have older versions consuming storage. Lifecycle rules can manage transitions and expiration, but locked versions cannot be expired before their Object Lock retention ends. [Versioning and lifecycle] [S3 Object Lock]
Think through these questions before enabling automatic deletion:
- How long must a reviewer be able to inspect a past capture?
- Do manifests and screenshots have the same retention period?
- Should the catalog record be removed when its image expires?
- Does Versioning require a separate noncurrent-version expiration rule?
- Do legal hold or Object Lock requirements apply?
7. Versioning, Object Lock, and recovery
Use Versioning for recovery
Versioning protects revisions under the same key and helps recover from accidental overwrites or deletions. It complements distinct timestamped keys: use unique keys to model capture history, then Versioning as a recovery layer. A delete marker can hide a key without immediately removing prior versions. Plan how to inspect, restore, and eventually expire those versions. [Retaining multiple object versions]
Use Object Lock only for immutability requirements
Object Lock provides WORM-style retention for a fixed period or indefinitely. It is for deliberate protection against deleting or overwriting retained object versions, not a routine substitute for Versioning or lifecycle management. A protected version cannot be removed by Lifecycle expiration before its retention ends; a delete marker may still be created. Object Lock also does not protect access to encryption keys or prevent KMS key deletion. [Object Lock]
Consider a second Region only when the recovery objective calls for it
Cross-Region Replication copies objects asynchronously and requires Versioning on both source and destination buckets. It can support a separate geographic copy, but adds operational choices and costs. Decide based on recovery time and recovery point needs, and verify current pricing and configuration requirements. [S3 replication]
8. Audit and inventory
Use CloudTrail data events or S3 server access logging when you need object-level access records. S3 Inventory can provide scheduled reports of objects and selected attributes, including encryption and Object Lock information. Inventory is not real-time: AWS says report generation can take up to 48 hours. For current Object Lock retention values, query the relevant object directly rather than treating Inventory as live state. [CloudTrail data events for S3] [S3 server access logging] [S3 Inventory]
9. Performance, reliability, and cost
- Capture cost and upload cost are separate. The browser work happens before S3 receives the image. S3 storage, requests, retrieval, replication, selected storage classes, and KMS usage can affect the archive bill. Estimate from capture volume, average file size, retention, read frequency, and current regional pricing; no project-specific estimate follows from the available facts.
- Control image size. Choose an appropriate viewport, output format, and image quality. Full-page captures can be substantially larger than viewport captures. If you resize or compress, preserve the original only when fidelity or audit requirements justify it.
- Keep uploads independent and retryable. Use a unique key for each run so a retry does not silently replace a prior record. If a retry should represent the same logical capture, keep a stable run ID and make the upload idempotent at the application layer.
- Write the image and manifest carefully. Two separate S3 puts are not a transaction. A reader may briefly see one without the other if the process fails midway. Write the image first, then the manifest as the commit marker, or record job state in a database and reconcile incomplete captures.
- Check integrity and status. Store a cryptographic checksum in the manifest and verify it when needed. Do not assume an HTTP success from the page capture means both S3 objects were written; handle SDK exceptions and log the bucket, key, request context, and capture ID without logging secrets.
- Know the storage durability claim’s scope. AWS describes S3 Standard as designed for 99.999999999% durability and 99.99% availability of objects over a given year. These service figures do not determine your application retention period or protect against every account, key-management, or regional recovery issue. [S3 data protection]
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
AccessDenied on upload |
The role lacks s3:PutObject, the bucket policy denies the request, or a KMS key policy blocks encryption. |
Check the caller identity, bucket and prefix policy, Region, and KMS permissions if using SSE-KMS. Keep permissions limited to the required archive prefix. |
| Bucket name or endpoint error | Wrong bucket, Region mismatch, or malformed endpoint configuration. | Confirm the bucket name and Region, then use the SDK’s normal regional endpoint handling. |
| Image exists but manifest is missing | The second upload failed after the image succeeded. | Treat the manifest as a commit marker; retry the manifest write or run a reconciliation task for orphan images. |
| History appears to have vanished after deletion | A delete marker hides the current key in a versioned bucket. | List object versions and restore or copy the intended prior version. Check lifecycle rules before restoring. |
| Storage remains high after expiration | Noncurrent versions remain, or delete markers have not been cleaned up. | Review Lifecycle rules for noncurrent versions and expired delete markers. Check whether Object Lock is preventing expiration. |
| Objects are unexpectedly public or unreadable | Public access settings, IAM/bucket policies, or account-level controls conflict. | Keep Block Public Access enabled and inspect effective policies and the identity used to read the object. |
| Screenshot is blank or incomplete | Navigation completed before client-rendered content, a resource failed, or the wait condition never matched page behavior. | Use an appropriate navigation condition, wait for a stable selector, and record status and final URL. Avoid an indefinite network-idle wait on pages with persistent connections. |
| Duplicate capture keys collide | Timestamp granularity is too coarse for concurrent captures. | Add a unique capture ID or random suffix to every key. |
| Inventory is missing recent changes | Inventory is scheduled and can take up to 48 hours to generate. | Use direct S3 API queries for current state; use Inventory for periodic reporting. |
11. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can return a screenshot in a single GET request; upload the returned bytes to your own S3 bucket and write the same manifest described above. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. After capture, upload the image bytes to S3 with the same unique-key and manifest approach.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
12. Frequently asked questions
Should I store a screenshot’s original page URL?
Yes, when useful for provenance and rediscovery, but consider whether the URL contains personal or secret query parameters. Redact sensitive values before writing the manifest.
Can S3 search my manifest JSON by page title?
Not as a general query over JSON object bodies. Keep an index in a database or catalog if you need search by title, hostname, tags, or time across a large archive.
Should screenshots and manifests use the same retention?
That depends on whether the manifest has independent audit or indexing value. Apply and verify lifecycle behavior for each prefix or object type.
Does Versioning create a screenshot timeline automatically?
No. It preserves revisions of the same key. A distinct key per capture plus an index or manifest is a clearer application-level timeline.
Can I share archived screenshots publicly?
Keep the archive private by default. If sharing is required, have an authorized application issue limited, short-lived access rather than exposing the whole bucket.


