How to Take a Website Screenshot at the Same Time Every Day with Puppeteer
Capture a website once a day with Puppeteer, save date-stamped screenshots, and schedule the script with cron or GitHub Actions in an explicit timezone.
Use Puppeteer to capture the page, then use a scheduler to run the script every day. Puppeteer takes the screenshot; it does not provide a durable daily schedule. For a reliable setup, make the target time zone explicit, save each run under a date-based filename, and run the script from an operating-system scheduler or hosted workflow that can launch it even when your development machine is off.
This guide uses Node.js with Puppeteer, then shows a GitHub Actions schedule and the decisions to make for cron or an in-process scheduler. The example captures a full-page PNG after navigation reaches networkidle2. That condition is only a starting point: dynamic pages may need a selector or other site-specific readiness check.
1. Create a Puppeteer screenshot script
Install Puppeteer in a Node.js project. Puppeteer currently requires Node.js 22.12 or later; check the official system requirements for supported browser platforms and Linux dependencies.
mkdir daily-screenshots
cd daily-screenshots
npm init -y
npm install puppeteer
mkdir -p screenshots
Save this as capture.mjs. Set TARGET_URL in the environment when you run it. The script creates its output directory, uses a stable viewport, gives navigation and capture a time limit, saves one PNG per UTC date, and closes the browser even if capture fails.
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
import puppeteer from 'puppeteer';
const targetUrl = process.env.TARGET_URL;
if (!targetUrl) throw new Error('Set TARGET_URL to the page to capture');
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 60_000);
const dateLabel = new Date().toISOString().slice(0, 10); // UTC date
const outputPath = path.join(outputDir, `${dateLabel}.png`);
await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.goto(targetUrl, {
waitUntil: 'networkidle2',
timeout: timeoutMs,
});
await page.screenshot({ path: outputPath, fullPage: true, type: 'png' });
console.log(`Saved ${outputPath}`);
} finally {
await browser.close();
}
The official Puppeteer screenshot guide uses Page.screenshot() and demonstrates navigation with networkidle2. See Puppeteer screenshots, Page.screenshot API, and ScreenshotOptions.
Run it once manually before scheduling:
TARGET_URL=https://example.com node capture.mjs
On Windows PowerShell, set the variable for the command like this:
$env:TARGET_URL = "https://example.com"
node .\capture.mjs
Make the filename match your intended day
toISOString() labels files in UTC. If the desired daily schedule follows another time zone, use that same IANA time zone to create the filename. Otherwise a run shortly after midnight local time could receive the previous or next UTC date. For example, use America/New_York or Europe/London as an explicit configuration value and format the date in that zone. Keep one date convention for the schedule and filenames.
The example overwrites the same day’s file if run more than once. If retries should preserve every attempt, include a timestamp in the name, or write to a temporary file and rename it only after the capture succeeds. For a single daily final file, a date-based path makes reruns idempotent.
2. Choose what the screenshot captures
Set the viewport before navigation so responsive layout is repeatable. The viewport is the browser’s CSS pixel size; deviceScaleFactor controls pixel density. A full-page capture is useful for an archive, but can create very tall images and increase memory and storage use.
| Need | Puppeteer option | Notes |
|---|---|---|
| Visible viewport only | fullPage: false |
This is the default. It captures the viewport, not the entire document. |
| Whole document | fullPage: true |
Can be very tall; lazy-loaded sections may require scrolling or a page-specific load routine first. |
| Specific rectangle | clip: { x, y, width, height } |
Use a fixed coordinate region when the page layout is stable. |
| PNG | type: 'png' |
Default image format; suitable when you want lossless output. |
| JPEG | type: 'jpeg', quality: 80 |
Smaller lossy images; quality is an integer from 0 to 100. |
| WebP | type: 'webp', quality: 80 |
Supported by the screenshot API where the installed browser supports it. |
Screenshot options and defaults can change with Puppeteer versions; consult its API reference when pinning or upgrading dependencies.
3. Wait for the page to be visually ready
A navigation event finishing does not mean every page looks settled. networkidle2 waits for a period with no more than two network connections, but analytics, long polling, delayed images, animation, or client-side updates can make it unsuitable. Choose a readiness signal based on the site:
- Known content element: after
goto, callawait page.waitForSelector('.report-ready', { timeout: timeoutMs })for a selector that appears when the content is usable. - Known application condition: use
page.waitForFunctionto wait until a specific DOM value or state is ready. - Fixed delay: use
await new Promise(resolve => setTimeout(resolve, 2000))only when the page has a predictable delay and no better signal. - Lazy images: scroll through the document before capturing if below-the-fold images load only when they approach the viewport; then wait for the relevant images to finish loading.
Do not use a long fixed sleep as a substitute for diagnosing a page that never becomes idle. A selector-based readiness check is usually more specific, and its timeout gives a clear failure when the expected content never appears.
4. Schedule the script to run daily
Pick a scheduler based on where the job must run and how much operational control you need.
| Scheduler | Process must stay running? | Time zone and restart behavior | Best fit |
|---|---|---|---|
| System cron | No; cron launches the command | Runs on the configured host. Check that host’s timezone and DST behavior; the machine must be on. | A server or computer you operate continuously. |
| GitHub Actions schedule | No local process; GitHub runs the workflow | Schedules default to UTC and support an IANA timezone. A scheduled time in a skipped spring-forward hour will not run at that local time. | A repository-based job whose files can be uploaded as workflow artifacts. |
| Node Schedule | Yes; the Node process must remain alive | In-process schedules are lost when the process exits or restarts; persistence is not provided by the scheduler. | An already-running Node service where an in-process timer is appropriate. |
The Node Schedule project recommends actual cron where persistence is required. See Node Schedule documentation. GitHub documents scheduled workflows, timezone configuration, and schedule behavior in Events that trigger workflows: schedule. Do not assume a scheduler starts at an exact second unless its documentation guarantees that.
Option A: GitHub Actions
Commit capture.mjs and package.json, then create .github/workflows/daily-screenshot.yml. This example runs at 09:00 in New York time, uploads the screenshot even if the capture step fails, and fails the workflow when capture fails. Replace the schedule and zone to match your requirement.
name: Daily website screenshot
on:
schedule:
- cron: '0 9 * * *'
timezone: America/New_York
workflow_dispatch:
jobs:
capture:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22.12'
- run: npm ci
- name: Capture page
run: node capture.mjs
env:
TARGET_URL: ${{ secrets.TARGET_URL }}
OUTPUT_DIR: screenshots
- name: Upload screenshot
if: always()
uses: actions/upload-artifact@v4
with:
name: daily-screenshot-${{ github.run_id }}
path: screenshots/
if-no-files-found: ignore
Set TARGET_URL in the repository’s Actions secrets. GitHub Actions scheduled runs can be delayed or dropped under some conditions, and scheduled workflows require the workflow file to exist on the default branch. Check the current GitHub documentation for limitations, timezone syntax, and skipped-hour behavior before relying on a local-time schedule.
Workflow artifacts are retained for a limited period according to repository settings. For a long-term archive, add a storage destination appropriate to your environment and credentials. Avoid committing screenshots to a public repository if the target page or image contains private information.
Option B: system cron
On a Linux or Unix host, install dependencies in a stable project directory and add a cron entry. This example launches at 09:00 according to the cron host’s configured timezone and appends stdout and stderr to a log:
0 9 * * * cd /opt/daily-screenshots && TARGET_URL='https://example.com' /usr/bin/node capture.mjs >> /var/log/daily-screenshot.log 2>&1
Use absolute paths in cron because its environment is minimal. Configure the host’s timezone deliberately, make sure the machine is on and connected, and check how its cron implementation handles daylight-saving transitions. Protect logs and screenshots if they may contain credentials or sensitive page data. Add log rotation or a retention policy so failure logs and output storage do not grow without bound.
Option C: Node Schedule in a running service
Node Schedule is an in-process scheduler. The process must stay up, and a restart loses the pending schedule until the application registers it again. Use a durable external scheduler when missed runs across process restarts matter.
import schedule from 'node-schedule';
schedule.scheduleJob('0 9 * * *', async () => {
// Run the capture logic here, or spawn capture.mjs and record its exit status.
console.log('Daily screenshot job started at', new Date().toISOString());
});
This cron expression does not by itself establish a time zone. Check the library version’s documentation for supported timezone and daylight-saving behavior, and keep a monitor or process manager in place for the long-running Node process.
5. Keep daily runs reliable
- Make runs observable: log the scheduled start, actual start, target URL (without secrets), output filename, and outcome. Send failures to a place someone checks.
- Use bounded timeouts: set navigation and selector timeouts so a hung page becomes a visible failure instead of a job that runs indefinitely.
- Handle retries deliberately: transient network failures may justify one or two retries with a short backoff. Avoid unlimited retries, which can create duplicate work and obscure persistent failures.
- Keep a consistent runtime: pin Node and Puppeteer versions through the lockfile. Browser and OS dependency changes can affect whether launch succeeds and how the page renders.
- Prevent overlapping runs: if a capture can last longer than a day or the scheduler retries while an earlier run is active, use a lock or scheduler-level concurrency control.
- Set retention: decide how long to keep screenshots and logs, then delete old files or configure artifact retention.
- Check output completeness: verify the file exists and has a nonzero size before marking the job successful.
Daily capture is not a guarantee that every run will start at the exact wall-clock instant. Scheduler load, host downtime, timezone rules, and daylight-saving changes can affect timing. If a capture must happen within a strict window, choose a platform with documented scheduling guarantees and alert on missed runs.
6. Performance, reliability, and cost
Self-hosted Puppeteer has no per-screenshot API fee, but it uses compute, browser memory, storage, and engineering time. Full-page shots, high device scale factors, large pages, and multiple concurrent browsers raise resource use. Reuse one browser for several captures in the same job if batching, but always close pages and the browser and cap concurrency.
A hosted workflow avoids maintaining an always-on machine, but execution quotas, artifact retention, and workflow scheduling policies depend on the provider and plan. System cron depends on the host staying available. An in-process timer depends on its Node process staying alive. Choose based on the cost of a missed capture, not only setup convenience.
For visual comparisons, use the same viewport, device scale, locale, authentication state, wait condition, and time zone every day. Dynamic ads, personalized content, rotating banners, and timestamps can create differences unrelated to a meaningful site change. Keep secrets out of source control and be mindful of privacy and site access rules.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Target closed or browser launch failure |
Unsupported Node version, missing Linux packages, or browser installation mismatch. | Use Node 22.12+, review Puppeteer’s system requirements for the host, and install the required OS libraries. |
| Navigation timeout | The page never reaches the chosen wait condition, often due to long-lived requests. | Try a more appropriate waitUntil such as domcontentloaded, then wait for a specific ready selector. Keep a bounded timeout. |
| Screenshot is blank or incomplete | Capture happened before client-side rendering or images finished loading. | Wait for a meaningful selector or app condition; scroll to trigger lazy images and wait for them before full-page capture. |
| Screenshot layout changes daily | Viewport, device scale, locale, page data, or personalization differs. | Fix viewport and browser settings, use a stable account/session when appropriate, and account for expected dynamic content. |
| Every run overwrites one file | The output filename is constant or date formatting is not what you expected. | Use a date-based filename and choose UTC or an explicit local IANA time zone consistently. |
| Job runs at the wrong hour | Scheduler uses UTC, host timezone differs, or daylight-saving rules changed. | Set an explicit timezone if supported, verify the scheduler’s DST policy, and inspect actual run timestamps. |
| It works locally but fails in cron or Actions | Scheduler environment has different paths, environment variables, permissions, or browser dependencies. | Use absolute paths, explicitly provide environment variables, inspect job logs, and match the runtime to Puppeteer’s requirements. |
| No scheduled workflow runs | Workflow is not on the default branch, schedule syntax or timezone is invalid, or the platform delays scheduled work. | Check the workflow’s Actions status and current schedule documentation; use manual dispatch to confirm the capture itself works. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, so a daily scheduler can call the endpoint instead of installing and maintaining a browser. The ScreenshotNeo API documentation lists the request options and response headers.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed; response headers indicate the page verdict and billing status. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. Learn more at ScreenshotNeo and read the API docs, then sign up for 1,000 free screenshots a month with no card.
Frequently asked questions
Does Puppeteer run the schedule itself?
No. Puppeteer controls the browser and captures the page. Use cron, a hosted workflow, or a running Node scheduler to launch the capture at the desired time.
Will a daily cron job run at the same local time after daylight-saving changes?
That depends on the scheduler and its timezone configuration. Set an IANA time zone where supported and check the scheduler’s documented spring-forward and fall-back behavior.
Should I use networkidle2 for every site?
No. It is one possible navigation condition. Pages with ongoing requests or delayed rendering may need a selector or application-specific readiness check.
How do I keep the screenshot files for months?
Choose a persistent storage destination and retention policy. Hosted workflow artifacts may expire based on repository settings, while local files remain only as long as the host and cleanup policy preserve them.


