How to Schedule Website Screenshots With an AI Agent
Build a recurring website screenshot workflow with Playwright and GitHub Actions, and learn where an AI agent helps, how to store captures, and how to troubleshoot missed runs.
To schedule website screenshots with an AI agent, divide the job into two parts: a browser capture task that waits for the page state you need and saves an image, and a scheduler that runs that task on a cadence. Use Playwright for repeatable captures and GitHub Actions for a straightforward recurring schedule. Add an AI agent when navigation or page interpretation varies; keep the schedule, capture settings, output naming, and verification deterministic.
This guide builds a recurring Playwright capture with GitHub Actions, explains the choices that affect screenshot consistency, and shows how an agent can fit into the workflow without making the whole process unpredictable.
1. Choose what the capture should represent
Before writing code, define the target URL and the exact page state you want to compare over time. Decide whether you need the visible viewport, the full scrollable page, or one element. Playwright supports all three, along with PNG, JPEG, and WebP output and CSS-pixel or device-pixel scaling. See the Playwright screenshot guide and Page API.
| Capture | Use it for | Trade-off |
|---|---|---|
| Viewport | A known screen area, dashboard panel, or above-the-fold layout | Does not include content below the visible browser window |
| Full page | A page-level visual record that includes scrollable content | Can produce a tall, large image and include content irrelevant to the change you are watching |
| Element | A stable component such as a pricing table or chart | Requires a selector that continues to identify the intended element |
For meaningful comparisons, keep the browser engine and version, viewport, device scale, capture scope, and output format consistent. Otherwise, a difference between files may come from capture setup rather than a site change. This is implementation guidance based on the capture controls exposed by Playwright.
2. Build the repeatable Playwright capture
The following Node.js script takes a URL from an environment variable, waits for a site-specific ready condition, and captures either the viewport, full page, or a selected element. It uses a fixed viewport and CSS-pixel scale to make runs more comparable. Choose a readiness selector that means the relevant content is actually ready; replace the example selector with one from your page.
// capture.mjs
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
const url = process.env.TARGET_URL;
if (!url) throw new Error('Set TARGET_URL to the page to capture');
const mode = process.env.CAPTURE_MODE ?? 'viewport'; // viewport | fullpage | element
const selector = process.env.CAPTURE_SELECTOR ?? 'main';
const readySelector = process.env.READY_SELECTOR ?? 'main';
const output = process.env.OUTPUT_FILE ?? 'artifacts/site.png';
await mkdir('artifacts', { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
colorScheme: 'light',
locale: 'en-US',
timezoneId: 'UTC',
});
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
if (!response || !response.ok()) {
throw new Error(`Navigation did not return a successful response: ${response?.status() ?? 'no response'}`);
}
// Prefer a page-specific readiness signal over an arbitrary sleep.
await page.locator(readySelector).waitFor({ state: 'visible', timeout: 30_000 });
// Optional: disable motion to reduce variation from animation timing.
await page.addStyleTag({ content: `*, *::before, *::after { animation: none !important; transition: none !important; caret-color: transparent !important; }` });
const options = { path: output, type: 'png', scale: 'css', animations: 'disabled' };
if (mode === 'fullpage') {
await page.screenshot({ ...options, fullPage: true });
} else if (mode === 'element') {
await page.locator(selector).screenshot(options);
} else if (mode === 'viewport') {
await page.screenshot(options);
} else {
throw new Error(`Unknown CAPTURE_MODE: ${mode}`);
}
console.log(`Saved ${output} (${mode}) for ${url}`);
} finally {
await browser.close();
}
Install and run it locally from a new project directory:
npm init -y
npm install playwright
npx playwright install chromium
TARGET_URL=https://example.com OUTPUT_FILE=artifacts/example.png node capture.mjs
To capture a full page, set CAPTURE_MODE=fullpage. To capture an element, set CAPTURE_MODE=element and CAPTURE_SELECTOR='.pricing-table'. Set READY_SELECTOR to a selector that only appears when the meaningful page content is ready. If the site updates data after rendering, wait for the updated value or application-specific state, not merely the initial document load.
3. Add an AI agent only where the page needs judgment
A fixed script is preferable for a stable URL and known selector. Use an AI computer-use agent if the workflow must interpret the current page, navigate a variable interface, or decide which visible control to use. The model proposes actions; your application still needs to execute those actions, manage browser state and access requests, and verify the result. OpenAI documents a computer-use session flow that includes handling access requests and reviewing completed activity in its computer-use documentation. Google Cloud also documents isolated browser sandboxes controllable through API requests or Chrome DevTools Protocol: Computer Use sandbox documentation.
A robust design keeps the agent’s authority narrow:
- Start from a known page or a restricted set of allowed domains.
- Give the agent a task such as navigating to a report and opening a named section, rather than broad instructions to explore.
- Let the agent perform only the UI actions that require interpretation; use ordinary Playwright locators for known controls where practical.
- After the agent finishes, assert that the expected page or element is present before taking the screenshot.
- Record whether the agent completed the task and retain enough run information to diagnose an unexpected capture.
Treat page content as untrusted input. A page can contain text that resembles instructions; it should not be allowed to expand the task or authorize unrelated actions. Use an isolated browser profile or sandbox, limit site access, and avoid giving a screenshot workflow credentials or permissions it does not need. An agent can interact with a browser, but that capability does not guarantee that it chose the right page state or produced a visually correct result.
4. Schedule the capture with GitHub Actions
Save the script as capture.mjs in the repository, then add .github/workflows/website-screenshot.yml. This example runs once daily at 09:17 UTC, installs Playwright and Chromium, captures the page, and uploads the image as a workflow artifact.
name: Scheduled website screenshot
on:
schedule:
- cron: '17 9 * * *'
workflow_dispatch:
permissions:
contents: read
jobs:
capture:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx playwright install --with-deps chromium
- name: Capture page
run: node capture.mjs
env:
TARGET_URL: https://example.com
READY_SELECTOR: main
CAPTURE_MODE: viewport
OUTPUT_FILE: artifacts/site-${{ github.run_id }}.png
- uses: actions/upload-artifact@v4
with:
name: website-screenshot-${{ github.run_id }}
path: artifacts/*.png
if-no-files-found: error
retention-days: 30
Commit the workflow to the repository’s default branch. GitHub Actions scheduled workflows use POSIX cron; GitHub’s current documentation says schedules run in UTC by default, can use an IANA timezone, run against the latest commit on the default branch, and have a shortest interval of five minutes. Scheduled workflows can be delayed under high load, especially at the start of an hour, and queued runs can be dropped. Scheduling at minute 17 avoids the most common top-of-hour bottleneck, but does not guarantee an exact start time. Read GitHub’s schedule event reference and workflow troubleshooting guide.
For a timezone-aware schedule, GitHub supports an optional IANA timezone field in the schedule entry. Check the current workflow syntax and expected daylight-saving behavior before relying on local wall-clock timing. If the time is operationally important, record the actual run time and alert on a missing or late capture.
5. Make the output useful over time
- Use unique filenames. Include a date, run ID, or timestamp so the next capture does not silently overwrite the previous one.
- Choose retention intentionally. Workflow artifacts are convenient for inspection; select a retention period appropriate to your comparison window and repository settings.
- Persist if you need a history. For long-term comparisons, upload captures to storage you control. Set retention and access permissions for that destination; the workflow example does not configure external storage.
- Keep run metadata. Store the URL, capture mode, viewport, browser version, timestamp, and completion status alongside the image. This helps distinguish site changes from configuration changes.
- Monitor the schedule itself. A successful workflow is not proof the correct page state was captured. Check that an image exists and, where useful, validate its dimensions or compare it with a known baseline.
6. Configure capture scope and consistency
Wait conditions
domcontentloaded waits for the document structure, not necessarily for client-rendered data, images, or a particular component. Add a locator wait for the content that matters. Use a fixed delay only for a known delay that cannot be represented by a better condition. Network idle can be a poor universal signal for pages with long polling or analytics traffic.
Viewport, full page, and element
The script sets the viewport to 1440 by 1000 CSS pixels. Change those values to match the view being monitored. A full-page image can be especially tall; if the target has infinite scrolling, loading more items as you scroll, or sticky elements, decide whether a full-page capture represents the intended state. Element captures fail if the selector matches no visible element; use a stable selector and wait for it first.
Format, scale, and motion
Playwright supports PNG, JPEG, and WebP. PNG is a practical lossless default; JPEG and WebP can reduce output size, with lossy quality settings available for those formats. CSS scale creates one image pixel per CSS pixel, while device scale captures device pixels and can create larger files on high-DPI settings. Disable animations when transient motion should not affect the capture, but consider leaving them enabled if the animation itself is what you are monitoring.
Authenticated or personalized pages
If the page requires a login, use a dedicated, least-privileged account and keep credentials in the scheduler’s secret store. Do not commit cookies, tokens, or storage-state files. Personalization, rotating banners, experiments, and locale settings can all change images between runs. Fix those inputs where possible, or expect them to appear as real visual differences.
7. Troubleshoot missed or incorrect captures
| Symptom | Likely cause | Fix |
|---|---|---|
| No scheduled workflow run | The workflow is not on the default branch, is disabled, or a public repository has been inactive | Confirm the workflow file exists on the default branch and is enabled. GitHub documents automatic disabling of scheduled workflows in public repositories after 60 days without activity; check the Actions settings and repository activity. |
| Run starts late or is absent during load | Scheduled events may be delayed or queued jobs dropped during high GitHub Actions load | Schedule away from the start of the hour, monitor run history, and use a scheduler with timing guarantees appropriate to a firm deadline. |
Timeout 30000ms exceeded waiting for selector |
The selector is wrong, appears only after interaction, or the page did not reach the expected state | Inspect the page and correct READY_SELECTOR; perform required navigation first; use a condition tied to the content rather than increasing timeouts without evidence. |
| Navigation returns an error or no response | Bad URL, network restriction, server failure, or a page that redirects unexpectedly | Check the URL and response status, confirm the runner can reach the host, and decide explicitly how redirects or non-2xx pages should be handled. |
| Image is blank or only a loading shell | The screenshot ran before client-rendered content or data appeared | Wait for a data-specific element or value. A successful document navigation alone does not prove the page is ready. |
| Element screenshot reports no element | The selector matches nothing, is in a different frame, or the element is hidden | Verify the selector, wait for it to become visible, and use a frame locator when the target is inside an iframe. |
| Images vary although the page looks unchanged | Animations, rotating content, viewport changes, locale, time, device scale, fonts, or personalization differ | Pin the browser and capture settings; disable irrelevant motion; set locale/timezone; account for genuinely dynamic regions in your comparison process. |
| Workflow cannot find the output | Output path differs from artifact upload path, or capture failed before writing a file | Keep OUTPUT_FILE under the uploaded directory, retain if-no-files-found: error, and inspect the earlier capture step’s logs. |
| Local run works but CI fails | Chromium dependencies are missing, secrets differ, or CI has a different network path | Install browsers with npx playwright install --with-deps chromium, configure required secrets, and inspect runner logs and target access rules. |
8. Performance, reliability, and cost
Each scheduled run starts a runner, installs or restores dependencies, launches a browser, loads the site, and writes an image. Browser startup and page load dominate many small capture jobs; avoid doing extra navigation or asking an agent to make decisions that a stable selector can handle. Cache dependencies only if it materially reduces setup time and keep browser versions aligned with the Playwright package.
Reliability has two separate parts: whether the scheduler starts the task, and whether the task captures the intended page. GitHub Actions is convenient for routine automation, but its schedule is not an exact-time guarantee. Add run monitoring and an artifact existence check. For strict deadlines, choose scheduling infrastructure with guarantees that meet the requirement. A browser-use agent adds flexibility but also introduces another variable step; validate its resulting page state before capturing.
Costs depend on the runner, browser or hosted-agent environment, storage, and how often captures run. The research sources do not establish a universal price or performance figure for this setup, so estimate using your own run duration, retention, and provider terms. Reduce unnecessary full-page images and retain only the comparison history you need.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. A single request captures a URL as an image or PDF, and its MCP tools let AI agents use screenshots directly. For a recurring workflow, your scheduler can call the API on its cadence, so you do not need to maintain a browser installation in the capture task.
Example cURL request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot; those steps can be turned off. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Can an AI agent run on a schedule by itself?
The agent needs an external trigger such as a scheduled workflow or job. The scheduler starts the task; the agent handles browser decisions within the task.
Should I use an AI agent for every screenshot?
No. For a fixed URL and known page state, scripted browser steps are usually easier to reproduce. Use an agent when the route or interaction genuinely needs interpretation.
Can GitHub Actions run screenshots more often than once a day?
Its current schedule documentation states a shortest interval of five minutes, though scheduled runs can be delayed or dropped under high load. Check the current GitHub rules before relying on a particular cadence.
Where should historical screenshots live?
Workflow artifacts work for short-term review. For a longer visual history, use storage with retention and access controls suited to your needs.


