How to Upload Recurring Website Screenshots to Amazon S3
Capture a website on a schedule with Playwright and GitHub Actions, then upload private screenshots to S3 using short-lived AWS credentials.
To upload recurring website screenshots to Amazon S3, run a browser capture on a schedule, authenticate to AWS, and upload each image to a deliberate S3 object key. A practical setup uses Playwright in GitHub Actions, GitHub OpenID Connect (OIDC) to assume a narrowly scoped AWS role, and the AWS CLI to copy the image. OIDC avoids storing long-lived AWS access keys as repository secrets.
This guide captures a public page. Site login, permission to automate capture, consent requirements, and whether you may retain the captured content are site-specific; check them before scheduling captures. Keep screenshots private unless you have deliberately chosen to publish them.
1. Choose what to capture and how to keep history
Decide what each image should show and whether each run should preserve history:
| Choice | Use it when | Trade-off |
|---|---|---|
| Viewport | You need the visible browser area. | Does not include content below the fold. |
| Full page | You need the complete scrollable page. | Can create a very tall image and take longer to render. |
| Element | You need one chart, panel, or component. | Requires a stable selector and the element to be present. |
| Timestamped S3 key | You need a history of captures. | Objects accumulate until you expire or transition them. |
| Stable S3 key | You only need the latest image. | Each run replaces the previous object. |
Playwright supports saving screenshots to a file or getting image bytes, full-page capture, and locator screenshots. The workflow below uses a full-page PNG and timestamped keys.
2. Create a minimal Playwright capture project
In a repository, create package.json:
{
"name": "scheduled-site-screenshot",
"private": true,
"type": "module",
"scripts": {
"capture": "node capture.mjs"
},
"dependencies": {
"playwright": "^1.55.0"
}
}
Install dependencies locally and commit the generated lockfile so the workflow can use npm ci:
npm install
npx playwright install --with-deps chromium
Create capture.mjs:
import { chromium } from 'playwright';
const targetUrl = process.env.TARGET_URL;
const outputPath = 'screenshot.png';
if (!targetUrl) {
throw new Error('Set TARGET_URL to the page to capture.');
}
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
const response = await page.goto(targetUrl, {
waitUntil: 'domcontentloaded',
timeout: 60_000,
});
if (!response) {
throw new Error('Navigation returned no main-document response.');
}
if (!response.ok()) {
throw new Error(`Navigation failed: HTTP ${response.status()} ${response.statusText()}`);
}
// Use a site-specific readiness signal when the page renders important data asynchronously.
await page.screenshot({ path: outputPath, fullPage: true, animations: 'disabled' });
console.log(`Saved ${outputPath} from ${targetUrl}`);
} finally {
await browser.close();
}
domcontentloaded avoids waiting for every resource, including analytics or long-lived requests. It does not guarantee that client-rendered content is ready. If the target fills the page asynchronously, wait for a meaningful selector, such as await page.locator('[data-capture-ready="true"]').waitFor({ state: 'visible', timeout: 30_000 }); before taking the screenshot. Replace that selector with one that actually exists on the site. For a simpler page where all resources must finish, use waitUntil: 'load'; do not assume that network idle is suitable for pages with persistent connections.
Capture one element instead
Replace the full-page screenshot line with a locator capture. Use the page’s real CSS selector:
const chart = page.locator('#revenue-chart');
await chart.waitFor({ state: 'visible', timeout: 30_000 });
await chart.screenshot({ path: outputPath });
For a viewport image, omit fullPage: true. Full-page images can be large; element capture is often a better artifact when the requirement is a specific component.
3. Configure AWS OIDC and a restricted S3 role
Configure GitHub as an OIDC identity provider in AWS and create a role that the intended GitHub repository can assume. Scope the role trust to the repository and, where appropriate, the branch or deployment environment. The precise trust policy depends on your repository and AWS account; use the subject claim for the intended context rather than a broad wildcard.
Grant the role only the access needed to upload under the chosen bucket prefix. The core object write permission is s3:PutObject on the destination object resource, for example arn:aws:s3:::YOUR_BUCKET/site-shots/*. Do not grant bucket-wide or account-wide permissions when a prefix is sufficient. Additional permissions may be needed if you choose features such as ACLs, object tags, Object Lock, or customer-managed KMS encryption; those depend on the bucket configuration.
GitHub’s workflow needs id-token: write to request an OIDC token. The AWS credentials action exchanges that identity for short-lived role credentials. Review and pin third-party Actions according to your supply-chain policy; GitHub identifies the AWS credentials action as third-party, not GitHub-certified.
4. Add a scheduled GitHub Actions workflow
Create .github/workflows/screenshot.yml. Replace the URL, role ARN, region, bucket, and prefix with your values. The example runs at 17 minutes past 06:00 UTC every day, away from the top of the hour.
name: Scheduled website screenshot
on:
schedule:
- cron: '17 6 * * *'
workflow_dispatch:
permissions:
contents: read
id-token: write
jobs:
capture-and-upload:
runs-on: ubuntu-latest
env:
TARGET_URL: https://example.com/
AWS_REGION: us-east-1
S3_BUCKET: YOUR_BUCKET
S3_PREFIX: site-shots/example-com
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
cache: npm
- name: Install dependencies
run: npm ci
- name: Install browser
run: npx playwright install --with-deps chromium
- name: Capture page
run: npm run capture
- name: Configure short-lived AWS credentials through OIDC
uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::123456789012:role/github-site-screenshot
aws-region: ${{ env.AWS_REGION }}
- name: Upload timestamped screenshot
shell: bash
run: |
timestamp="$(date -u +'%Y-%m-%dT%H-%M-%SZ')"
aws s3 cp ./screenshot.png "s3://${S3_BUCKET}/${S3_PREFIX}/${timestamp}.png" \
--content-type image/png
The workflow file must be present on the repository’s default branch for scheduled runs. GitHub schedules use UTC by default and can also support an IANA timezone. Local-time schedules can shift around daylight-saving transitions, so choose UTC when a fixed operational time is preferable. Scheduled events can be delayed or dropped during high load; they are not a precise monitoring clock. If you use this for alerting or exact capture intervals, choose a scheduler with guarantees that fit that requirement.
The sample action references are readable version tags. Pin actions to reviewed commit SHAs if required by your repository’s security policy. Keep the capture code’s dependencies locked using the committed lockfile.
5. Pick an S3 key and retention policy
The example writes a UTC timestamp into the key, preserving one object per successful run. To retain only the latest capture, upload to a stable key instead:
aws s3 cp ./screenshot.png "s3://${S3_BUCKET}/${S3_PREFIX}/latest.png" \
--content-type image/png
A stable key replaces the current object at that key. If bucket versioning is enabled, previous versions may still be retained and incur storage costs; understand the bucket’s versioning and expiration behavior before assuming an overwrite removes old data.
Use an S3 Lifecycle rule on the screenshot prefix to transition older captures to a lower-cost storage class or expire them when they are no longer needed. Lifecycle rules apply to existing and new matching objects, but they do not act as an exact real-time deletion timer. With versioning, configure the treatment of noncurrent versions as well as current objects. Set retention deliberately and confirm it matches any audit or privacy requirements.
S3 objects are private by default and new objects are encrypted by default. Encryption does not replace access control. Avoid putting credentials or sensitive customer details in object keys. Do not make a bucket public merely to simplify viewing screenshots; grant access through your normal private access path.
6. Check the result and diagnose failures
Run the workflow manually once with workflow_dispatch, then inspect the Actions run and verify the object key in the intended bucket and prefix. Check both the browser step and upload step: a successful upload can still contain an error page or an incomplete client-rendered view.
| Symptom | Likely cause | Fix |
|---|---|---|
| No scheduled run | Workflow is not on the default branch, cron is misread as local time, or GitHub delayed/dropped a queued event. | Confirm the file is on the default branch, interpret cron in UTC unless a timezone is configured, and avoid scheduling at minute zero. Use a separate monitoring scheduler when exact intervals matter. |
| AWS reports access denied | OIDC trust subject does not match the repository context, the role lacks s3:PutObject for the key, or an encryption/tag/ACL setting needs additional permissions. |
Check the role trust subject and exact bucket-prefix object ARN. Add only permissions required by the selected bucket features. |
| Could not assume role or no OIDC token | The workflow lacks id-token: write, the provider or trust relationship is misconfigured, or the role ARN is wrong. |
Set the workflow permission, verify the AWS OIDC provider and role trust, and check the account ID and role name. |
| Browser install or launch fails | Playwright’s browser binaries or operating-system dependencies are missing. | Run npx playwright install --with-deps chromium in the job after installing the package. |
| Screenshot is blank, partial, or missing data | The page has not finished rendering, requires a session, blocks automation, or the chosen selector does not match. | Check the run’s navigation status, wait for a site-specific ready selector, verify the selector and viewport, and handle permitted authentication explicitly. Respect the site’s access rules. |
| Navigation timeout | The site is slow, a resource hangs, or the timeout is too short for its main document. | Inspect the page and network behavior, select an appropriate navigation state, and adjust the timeout deliberately. Avoid waiting for every request if the page keeps connections open. |
| Upload succeeds but object is hard to find | The timestamped key differs from the expected location or the workflow used another bucket/region/prefix. | Print the non-sensitive destination bucket and key in the job log, then inspect that exact prefix and region. |
| Older screenshots still consume storage | Timestamped keys accumulate, bucket versioning preserves old versions, or lifecycle expiration has not run yet. | Review current and noncurrent version lifecycle rules and storage class transitions; lifecycle processing is not immediate. |
7. Performance, reliability, and cost
Browser startup, navigation, page rendering, and image upload all contribute to each run. Reusing a browser for multiple URLs can avoid repeated launches, but each page still needs its own readiness condition and error handling. Full-page captures may use more memory and produce larger files than viewport or element captures. Keep image dimensions and capture frequency aligned with the actual review or audit need.
Scheduled GitHub Actions is convenient for periodic capture, but its schedule can be delayed or dropped under load. Treat it as a best-effort trigger rather than a monitoring SLA. For audit history, timestamped keys make missed runs visible as gaps; for a dashboard that only needs the latest image, a stable key simplifies retrieval.
There is no meaningful universal S3 cost estimate without the AWS region, image size, capture frequency, retention duration, request volume, and storage class. Estimate storage for the number and size of retained objects, plus PUT requests and any retrieval or transition charges. Use the current AWS pricing page or calculator with your region and lifecycle design. Reduce costs by retaining only the history you need, applying lifecycle expiration or transition rules, and avoiding repeated captures of unchanged pages if your workflow can safely determine that no new image is needed.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL as PNG, JPEG, WebP, or PDF; the response includes headers that identify the page verdict and whether it was billed. Its capture flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server exposes screenshot, page-info, and PDF-capture tools to Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for request options and account setup. This one-call example saves a WebP response:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/ \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
To put an API result in S3, save the response to a file and upload it with the same AWS CLI command or SDK approach used above; ScreenshotNeo does not replace your S3 bucket, schedule, access policy, or retention decisions. ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
FAQ
Can I schedule screenshots more than once a day?
Yes. Add cron entries or choose a suitable interval, accounting for GitHub’s scheduling delays and the need to avoid heavy schedules at the top of an hour.
Can I store screenshots as JPEG or WebP?
Playwright’s screenshot API supports image output formats. Choose the format and file extension consistently, and set the matching S3 content type on upload. PNG is a straightforward default for screenshots with text and sharp edges.
Will screenshots be publicly accessible in S3?
No. S3 resources are private by default. Keep that setting unless public delivery is an explicit requirement, and remember that default encryption does not grant or restrict user access.
What if the page needs authentication?
Use only an authorized account and a site-approved capture method. Authentication setup depends on the site; avoid placing passwords or tokens in source code, logs, filenames, or public object paths.


