How to Use Puppeteer to Capture Product Screenshots on a Schedule
Schedule repeatable product screenshots with Puppeteer, from page readiness and capture options to a reliable GitHub Actions workflow.
Use Puppeteer’s Page.screenshot() in a Node.js script, then run that script on a cron schedule. Set a consistent viewport, choose a readiness condition that matches the page, write each capture to a predictable path, and close the browser in a finally block. Puppeteer handles browser automation and capture; your scheduler provides repetition.
This guide uses a local script and GitHub Actions as an example. The schedule and artifact configuration are easy to adapt to another runner. Puppeteer documents both page screenshots and element screenshots in its screenshot guide.
1. Set up the project
Use a supported Node.js version for your environment, then install Puppeteer. The puppeteer package downloads a compatible browser during installation; if your environment manages Chrome separately, use the corresponding Puppeteer configuration and executable path.
mkdir scheduled-product-shots
cd scheduled-product-shots
npm init -y
npm install puppeteer
Create a URL list named targets.json. Keep a stable name per target because it becomes part of the output filename.
[
{ "name": "product-home", "url": "https://example.com/" },
{ "name": "pricing", "url": "https://example.com/pricing" }
]
Validate target URLs and names before running an unattended job. Avoid names containing slashes or characters that are awkward in file paths.
2. Write the capture script
Save this as capture.mjs. It launches one browser for the run, opens a fresh page per target, uses a fixed viewport, waits for a product-specific selector, and saves a full-page PNG. Replace main with a selector that appears only after the useful product content is rendered. If a selector is not reliable for your target, use another readiness option in the next section.
import puppeteer from 'puppeteer';
import { mkdir, readFile } from 'node:fs/promises';
import path from 'node:path';
const outputDir = path.resolve('screenshots');
const targets = JSON.parse(await readFile('targets.json', 'utf8'));
await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
try {
for (const target of targets) {
if (!target.name || !/^https?:\/\//i.test(target.url)) {
throw new Error(`Invalid target: ${JSON.stringify(target)}`);
}
const page = await browser.newPage();
try {
await page.setViewport({ width: 1440, height: 1000, deviceScaleFactor: 1 });
page.setDefaultNavigationTimeout(45_000);
await page.goto(target.url, { waitUntil: 'networkidle2', timeout: 45_000 });
// Prefer a real content-ready signal when the product page has one.
await page.waitForSelector('main', { timeout: 15_000 });
const file = path.join(outputDir, `${target.name}.png`);
await page.screenshot({ path: file, type: 'png', fullPage: true });
console.log(`Captured ${target.url} -> ${file}`);
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
Run it locally with node capture.mjs. The screenshot path is resolved from the process working directory, so run the command from the project directory or use an absolute path. Supplying path writes the image to disk; without it, page.screenshot() returns image data instead. See the official Page.screenshot() reference and ScreenshotOptions reference.
3. Choose when a page is ready
The waitUntil option on page.goto() controls the navigation milestone Puppeteer waits for. It does not prove that a product’s client-side data, animations, or lazy images are ready.
| Condition | Useful when | Trade-off |
|---|---|---|
domcontentloaded |
The needed markup is available after the document is parsed. | Images, styles, and later application work may still be in progress. |
load |
Ordinary page resources should finish loading. | Some pages continue making requests or render important content afterward. |
networkidle2 |
The page settles while allowing some active requests. | Persistent traffic can still delay or prevent the condition. |
networkidle0 |
The page is expected to become fully quiet. | Analytics, polling, or websockets may keep it from becoming quiet before timeout. |
For a product page, a known content selector is often a stronger signal than network quiet alone. Combine a practical navigation milestone with waitForSelector() or waitForFunction() for a page-specific ready state. Set timeouts deliberately: a strict timeout exposes stalled pages, while a very long timeout makes a scheduled run slow to fail.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.waitForSelector('[data-testid="product-gallery"] img', { timeout: 20_000 });
Lazy-loaded images may not load until scrolled into view. For a full-page capture that depends on them, scroll through the document before taking the screenshot, then allow image requests to settle. This is page-specific: long pages and sites with infinite scrolling need a bounded policy rather than an unending scroll loop.
4. Pick the capture area and output
- Viewport: omit
fullPageor set it tofalsefor a consistent screen-sized image. - Full page: set
fullPage: trueto capture the document beyond the viewport. Very tall pages can take longer and produce large files. - One element: wait for the element and call its
screenshot()method. Puppeteer scrolls a hidden element into view before capturing it. - Clip: use
clipfor a specific rectangle when you need a fixed region. Check the current API reference for clip coordinates and capture behavior. - Format: the documented default is PNG. Set
typeto a supported format.qualityapplies to formats that support it and does not affect PNG. - Transparent background:
omitBackground: truehides the default white background where transparency is supported. - File path: set
pathexplicitly. A relative path uses the process working directory.
// A focused product card
const card = await page.waitForSelector('.product-card');
await card.screenshot({ path: 'screenshots/product-card.png', type: 'png' });
For repeat comparisons, keep viewport dimensions, device scale factor, locale, browser version, and page state stable. A fixed setup improves comparability but does not guarantee identical pixels across browser versions, operating systems, fonts, or changing page content.
5. Run it on a schedule with GitHub Actions
Create .github/workflows/product-shots.yml. This example runs daily at 08:00 UTC and also supports manual runs. It uploads the output as a workflow artifact for review or download. GitHub cron expressions use UTC; check the current GitHub Actions documentation for schedule behavior, runner limits, and artifact retention before relying on them.
name: Product screenshots
on:
workflow_dispatch:
schedule:
- cron: '0 8 * * *'
jobs:
capture:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: node capture.mjs
- uses: actions/upload-artifact@v4
with:
name: product-screenshots
path: screenshots/
if-no-files-found: error
retention-days: 14
Commit package-lock.json so npm ci installs the locked dependency tree. The example’s retention setting controls how long that artifact is retained under the action’s configuration; choose a period that meets your review and storage needs. If you want images in the repository instead, add a deliberate commit-and-push step and decide how to handle concurrent runs and generated diffs.
A third-party GitHub Marketplace action also demonstrates a scheduled screenshot workflow with its own options for URL lists, concurrency, retries, timeouts, output directories, and pull requests. Those options describe that action specifically and are not guarantees about GitHub Actions generally. Review its maintenance and permissions before adopting it.
6. Make scheduled captures reliable
- Use stable filenames. Derive filenames from controlled names rather than page titles, which can change or contain path characters.
- Bound each operation. Set navigation and selector timeouts, and set a job-level timeout so a hung browser cannot occupy a runner indefinitely.
- Close resources. Use
finallyfor browser and page cleanup so one failed URL does not leak the browser process. - Log target and outcome. Include the URL and output path in success logs. On failure, preserve the error message in the job output.
- Retry selectively. A small retry count can help with transient network failures. Do not retry invalid URLs, missing selectors caused by a changed page, or access-denied responses indefinitely.
- Keep the schedule explicit. Document the cron timezone and cadence. Avoid overlapping jobs if two runs could write the same files or commit competing changes.
- Retain useful history. Keep artifacts long enough for the intended comparison, and remove old files if you store captures in a repository.
For visual comparisons, record the capture time and conditions alongside each image. Product content can change between runs, and font or browser changes can alter pixels even when the layout is healthy.
7. Performance, reliability, and cost
Runtime grows with the number of URLs, page complexity, readiness waits, and full-page height. Reuse one browser process for a small batch, as in the example, while giving each URL its own page. Keep concurrency modest on constrained runners; too many simultaneous pages can increase memory use and cause slower or less reliable captures.
Set the viewport and image format based on the downstream use. A viewport capture is usually smaller and faster than a very tall full-page image. PNG preserves lossless detail but can be large; a supported lossy format with a chosen quality can reduce file size. Quality has no effect on PNG. Scheduler cost depends on the provider, runner size, cadence, and runtime; the supplied research does not establish provider pricing, so estimate from your own workflow usage and current provider terms.
Scheduled automation is not a guarantee that every target will be reachable at every run. Sites can change selectors, block automated browsing, serve region-dependent content, or have transient failures. Treat missing output as a failed job, alert on repeated failures, and make a manual workflow run available for diagnosis.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation times out | The page keeps requests open, is slow, or is unreachable from the runner. | Choose a suitable navigation milestone such as domcontentloaded, then wait for a meaningful selector. Keep a finite timeout and inspect the target’s availability. |
| Screenshot is blank or incomplete | The app has not rendered its main content, or lazy content has not loaded. | Wait for the content selector or app state; scroll boundedly if lazy images matter; capture after the relevant state appears. |
waitForSelector times out |
The selector changed, is wrong for this route, or content requires login or interaction. | Inspect the page structure and use a stable selector. Handle authentication only through a secure, intended test setup. |
| Images are missing in a full-page shot | Images load lazily as they approach the viewport. | Scroll the page in bounded steps, wait for image loading to settle, and then capture. Infinite-scroll pages need an explicit maximum. |
| Browser fails to launch in CI | Browser installation, system dependencies, or runner configuration is incomplete. | Install dependencies from the lockfile and follow Puppeteer’s current CI and configuration guidance for the selected runner. |
| File is not found after a successful run | The relative output path resolved from a different working directory, or the artifact step targets another directory. | Resolve output paths explicitly and make the workflow artifact path match the script’s output directory. |
| Captures differ between runs | Viewport, browser, fonts, locale, content, or timing changed. | Pin the environment where practical, use a stable ready signal, and compare only after accounting for expected page changes. |
| Scheduled job does not run when expected | Cron timezone assumptions or provider schedule behavior differ from expectations. | Use the provider’s documented timezone semantics, verify the expression, and keep manual dispatch for diagnosis. |
Or skip the browser setup
With ScreenshotNeo, one GET request returns a screenshot or PDF. See the API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, no card required.
FAQ
Can I capture only a product image or card?
Yes. Wait for the target element and use its screenshot() method; Puppeteer scrolls it into view if needed.
Does Puppeteer itself schedule the capture?
No. Puppeteer automates the browser and takes the image. A cron service or scheduled workflow starts the script at the chosen times.
Should I use full-page screenshots for visual regression?
Use them when changes anywhere on the page matter. For a stable comparison region, a fixed viewport or selected element reduces unrelated differences.
Can I save the image as data instead of a file?
Yes. Omit path and await the screenshot result, which can then be stored or uploaded by your script.


