How to Set Up a Monthly Website Screenshot Archive with Playwright
Build a recurring Playwright screenshot archive, schedule it with GitHub Actions, and choose storage that preserves the history you need.
A monthly website screenshot archive needs three parts: a Playwright script that captures each configured URL, a scheduler that runs it monthly, and storage whose retention matches how long you need the screenshots. The example below saves full-page PNGs in date-based folders and runs from GitHub Actions. Workflow artifacts are convenient for short-term retrieval, but GitHub’s default retention is 90 days, so choose durable storage for a multi-year archive.
1. Create the Playwright capture script
Set up a Node.js project, install Playwright, and install its Chromium browser using the official Playwright setup guide. Save this script as capture.mjs in the project root. Replace the sample URL list with the pages you want to preserve.
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
const urls = [
'https://example.com',
// Add one URL per page to archive.
];
const now = new Date();
const year = String(now.getUTCFullYear());
const month = String(now.getUTCMonth() + 1).padStart(2, '0');
const day = String(now.getUTCDate()).padStart(2, '0');
const browser = await chromium.launch();
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
});
for (const url of urls) {
const safeName = new URL(url).hostname.replaceAll('.', '-');
const dir = `archive/${year}/${month}`;
await mkdir(dir, { recursive: true });
await page.goto(url, { waitUntil: 'networkidle', timeout: 60_000 });
await page.screenshot({
path: `${dir}/${safeName}-${year}-${month}-${day}.png`,
fullPage: true,
animations: 'disabled',
});
}
} finally {
await browser.close();
}
The path groups captures by UTC year and month, and includes the date in each filename. If you run this more than once per day and want to keep every run, add a UTC time component to the filename; otherwise a later capture on the same date overwrites the earlier file.
Choose a useful page-ready condition
networkidle waits for network activity to settle. Some sites keep connections open or continuously fetch content, making that condition unsuitable. In that case, navigate with domcontentloaded or load, then wait for a page-specific selector:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.locator('main').waitFor({ state: 'visible', timeout: 30_000 });
Pick a selector that indicates the content you actually want to archive. A generic navigation event can finish before client-rendered content appears. For pages with lazy-loaded images, full-page capture may not itself cause every image to load; add a site-specific scroll or readiness step where needed, then verify the resulting archive files.
Capture viewport or full page
page.screenshot({ path }) captures the current viewport. Set fullPage: true to capture the full scrollable page. Full-page images can be very tall and large, so use viewport captures if the goal is to record the initial visible state or keep files smaller. Playwright’s screenshot guide and Page API reference describe screenshot options, including element capture, scale, animations, and output format.
Other useful options include scale: 'css' for one output pixel per CSS pixel or scale: 'device' for device-pixel scaling. To capture a single element, use page.locator('selector').screenshot({ path }). PNG is lossless and useful for visual comparisons; JPEG or WebP can reduce storage size when your workflow and review process support them. If you change browser, operating system, viewport, scale, or other rendering settings, record that change in your archive manifest because screenshots may differ.
2. Schedule a monthly run with GitHub Actions
Create .github/workflows/monthly-archive.yml. This schedule runs at 04:17 UTC on the first day of each month; the minute is offset from the start of the hour because GitHub notes scheduled runs can be delayed during periods of high Actions load. Scheduled workflows use UTC by default and run from the latest commit on the default branch. GitHub may delay or drop a scheduled run during high load, so use workflow_dispatch for manual runs and monitor whether monthly captures arrive.
name: Monthly website archive
on:
schedule:
- cron: '17 4 1 * *'
workflow_dispatch:
jobs:
capture:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: node capture.mjs
- uses: actions/upload-artifact@v5
with:
name: website-screenshots-${{ github.run_id }}
path: archive/
retention-days: 30
This workflow uses illustrative action version tags; check the current releases and your repository’s action policy when publishing or adopting it. The 30-day artifact setting is illustrative and shorter than GitHub’s documented 90-day default. Set retention deliberately, within your repository and account policies. The schedule only runs when the workflow file is on the default branch. See GitHub’s schedule event documentation and artifact documentation.
3. Choose archive storage and make it browsable
| Storage approach | Useful for | Trade-off |
|---|---|---|
| Workflow artifacts | Reviewing or downloading recent run output | Artifacts expire according to configured retention and account or repository policy; default retention is 90 days. |
| Repository files | A small archive that benefits from version history and simple browsing | Images grow repository size over time and may make clones and history expensive. |
| Object storage | A longer-lived archive with independent lifecycle and access settings | Requires choosing a provider, access policy, lifecycle, and cost model for your volume and retention needs. |
For long-term retention, upload the generated files to durable storage or commit them only when repository growth is acceptable. Do not treat ordinary workflow artifacts as permanent storage. Before choosing a destination, compare retention duration, access control, browsing and download workflow, portability, and expected storage volume. No storage price estimate is included because it depends on your requirements and provider.
Add a manifest alongside the images so future readers can interpret them. Useful fields include source URL, capture timestamp, viewport, browser and version, output format, and file path. A simple index or generated HTML page can link each URL’s captures in chronological order.
4. Keep captures useful and comparable
A screenshot records one rendered state at one time; it does not show everything a site did during the month or prove that the site was available continuously. Include the UTC capture time and environment in the manifest. If you plan to compare screenshots, keep the operating system, browser version, viewport, browser settings, and runner conditions consistent. Playwright identifies operating system, browser version, settings, hardware, power source, and headless mode as factors that can change visual output. Read its visual comparison guidance before interpreting small differences as site changes.
Images and related reports can contain private page content. Restrict access to the repository or storage location as appropriate. Playwright advises uploading reports and traces only to trusted artifact stores, or encrypting them before upload; apply the same care to screenshots that may expose account data or internal pages. See the Playwright CI guide.
5. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The script times out waiting for navigation | The page never reaches the selected load state, often because it maintains network activity. | Use domcontentloaded or load, then wait for a meaningful selector. Raise the timeout only when the page genuinely needs more time. |
| The screenshot is blank or missing client-rendered content | Capture happened before the page’s own content was ready. | Wait for a page-specific visible selector or another site-specific readiness condition before capturing. |
| Some images are absent in a full-page capture | Lazy-loaded content may not have loaded before capture. | Add scrolling or another site-specific loading step, then wait for the relevant images or content before taking the screenshot. |
| A later run replaced an earlier image | The filename uses only the date, so multiple daily runs share a path. | Add hours, minutes, and seconds (UTC) or a unique run identifier to the filename. |
| No monthly workflow run appears | The schedule runs from the default branch and can be delayed or dropped under high load. | Confirm the workflow is on the default branch, check Actions status, use manual dispatch to diagnose, and schedule away from minute zero. |
| Old screenshots are no longer available | Workflow artifacts expired under their retention policy. | Increase allowed retention or copy each run to durable storage with a lifecycle policy that matches the archive’s required lifetime. |
| Visual diffs change even when the page seems unchanged | Browser or runner environment changed, or the page includes dynamic content. | Standardize the environment and capture settings; record them in the manifest and account for site content that changes between runs. |
| Workflow cannot install Chromium dependencies | The runner or install command does not match the project’s Playwright setup. | Use the Playwright browser installation command for the installed package and an environment supported by its CI guidance. |
6. Performance, reliability, and cost considerations
Monthly cadence keeps run volume modest, but each URL adds navigation and image storage. Capture only pages that answer a defined monitoring or historical question. Full-page shots usually cost more time and storage than viewport shots; choose image format and dimensions according to whether you need pixel-level comparison or compact records. Keep a small URL list initially, then estimate archive growth from the number of URLs, captures, and average image size before committing images to a repository.
A hosted scheduled workflow is convenient but not a guarantee of an exact run time. GitHub documents possible schedule delays and missed runs under load. If missing a month’s record would matter, check workflow outcomes and alert on failures through your normal operational process, or use a scheduler and storage system whose guarantees meet that need. The capture itself can fail because of site availability, navigation, bot checks, or content changes; retain enough run logs to identify which URL failed without exposing secrets.
GitHub artifact retention may be configurable, but ordinary artifacts expire. Durable storage introduces provider-specific storage, request, and retrieval costs; estimate those from expected image volume and retention rather than assuming a universal price. Protect credentials if the archive includes authenticated pages, and avoid putting secrets directly in source control or logs.
Or skip the browser setup
If you want a screenshot archive without maintaining a browser install and capture script, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. You still need a scheduler and a storage destination for a monthly archive, but the capture call is a single request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot tools. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. These features are available on every plan. To schedule an archive, call the API from your existing monthly job and save each result under a dated path.
Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card required.
FAQ
Does a monthly archive need to use full-page screenshots?
No. Use full-page capture when the entire document matters; use viewport capture when you want a stable record of what a visitor first sees or need smaller files.
Can I run the capture locally before enabling the schedule?
Yes. Run node capture.mjs from the project root after installing the package and Chromium, then inspect the output paths and images before relying on the scheduled workflow.
Will screenshots prove that the site stayed available all month?
No. Each image records a single successful rendered capture attempt. It is a periodic visual record, not continuous availability monitoring.


