ScreenshotNeo

BlogComparisons

Best Open-Source Tools for Scheduled Chromium Website Screenshots

Compare shot-scraper, Playwright, and Puppeteer, then build a recurring Chromium screenshot workflow with GitHub Actions, local scheduling, and practical reliability guidance.

By the ScreenshotNeo team4 October 20269 min read

For a recurring screenshot archive, start with shot-scraper and run it from GitHub Actions or another scheduler. Choose Playwright when you need custom browser scripting or screenshot-based visual comparisons. Puppeteer is a practical alternative when its Chromium-oriented browser automation API fits your workflow. These recommendations follow documented capabilities; they are not based on controlled speed or reliability benchmarks.

A screenshot tool captures the page. A scheduler starts the capture job at an interval. Plan both pieces, plus where the resulting images go. For scheduled CI captures, remember that a cron schedule is not a precision timer: GitHub documents possible delays and dropped queued runs under high load.

1. Choose a tool for the capture job

Tool Good fit What its documentation supports Trade-offs
shot-scraper Repeatable command-line archives and configured multi-page captures URL screenshots, element selection, multi-capture configuration, JavaScript, and a GitHub Actions workflow example. It captures pages; a workflow or other scheduler must run it. Its example commits captures to a repository, which is only one possible storage choice.
Playwright Custom browser automation, full-page or element captures, and visual checks Page screenshots, full-page screenshots, image-buffer output, locator screenshots, and screenshot assertions through Playwright Test. Visual comparisons depend on the browser and runner environment. Keep the environment consistent with the baseline.
Puppeteer Chromium-oriented browser automation and element captures Page and selected-element screenshots. Puppeteer documents that it attempts to scroll a hidden element into view for a screenshot. The reviewed documentation does not establish that it is faster, more reliable, or universally better than Playwright.

Compare your options against five needs:

  1. Capture target: viewport, full scrollable page, or a selected element.
  2. Page state: interactions, authentication, JavaScript, and waits required before capture.
  3. Schedule: local cron, CI workflow, or another recurring runner, and whether its timing is suitable.
  4. Purpose: archive images, or also compare them against reference screenshots and review changes.
  5. Operations: stable browser and runner environment, output storage, and a plan for delayed or missed runs.

2. Build a recurring archive with shot-scraper

shot-scraper is a good first choice when the job is a configured list of URLs captured on a recurring schedule. Its documentation includes a GitHub Actions example that installs the package and browser dependency, captures entries defined in shots.yml, and commits outputs. The following is a minimal workflow pattern based on that documented approach; check the current shot-scraper manual for the exact configuration options you need.

Define the URLs

# shots.yml
- url: https://example.com/
  output: shots/example-home.png
- url: https://example.org/
  output: shots/example-org.png

Add a scheduled workflow

# .github/workflows/screenshots.yml
name: Scheduled screenshots

on:
  workflow_dispatch:
  schedule:
    - cron: "17 8 * * *"

permissions:
  contents: write

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.x"
      - name: Install shot-scraper and browser
        run: |
          pip install shot-scraper
          shot-scraper install
      - name: Capture configured pages
        run: shot-scraper multi shots.yml
      - name: Save updated captures
        run: |
          git config user.name github-actions[bot]
          git config user.email 41898282+github-actions[bot]@users.noreply.github.com
          git add shots/
          git diff --cached --quiet || git commit -m "Update scheduled screenshots"
          git push

The cron expression above requests a run daily at 08:17 UTC. GitHub Actions schedules use cron syntax, default to UTC, run on the default branch, and support a shortest interval of five minutes. GitHub warns that high load can delay runs or cause queued jobs to be dropped, so avoid scheduling at a precise business deadline. See GitHub’s workflow event documentation for current schedule behavior.

The commit-and-push step is appropriate only if image files belong in the repository. For a large or long-lived archive, consider a separate storage destination or retention policy; choose one that fits your infrastructure. Keep credentials out of the YAML: store secrets in the runner’s secret mechanism and expose them only to the capture step that needs them.

3. Use Playwright when capture needs browser logic

Playwright suits jobs where a simple URL list is not enough: you may need to navigate, wait for a specific state, capture a full page, or feed a screenshot into a comparison workflow. This runnable Node.js script saves a full-page PNG. Install Playwright and its Chromium browser in the same environment used by your scheduler.

// capture.mjs
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto('https://example.com/', { waitUntil: 'networkidle', timeout: 60_000 });
  await page.screenshot({ path: 'example-full.png', fullPage: true });
} finally {
  await browser.close();
}

Run the script from a recurring runner, for example with a scheduled CI workflow or the operating system’s cron service. Playwright documents viewport, full-page, buffer, and locator screenshot methods in its screenshot guide. For a selected element, locate it and call locator.screenshot({ path: 'element.png' }). For a visual assertion, use Playwright Test’s screenshot comparison support rather than treating a saved image as an automatic pass/fail check.

networkidle can be a poor fit for pages with long-lived network connections or continuous background requests. In that case, wait for a meaningful selector or application state, with a bounded timeout, then capture. A page that has finished network activity is not necessarily visually ready, and a page may be ready while analytics requests continue.

4. Use Puppeteer for Chromium page or element captures

Puppeteer is another option when you want a Chromium-focused automation script. This example opens a page and saves its full-page screenshot:

// capture-puppeteer.mjs
import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1440, height: 900 });
  await page.goto('https://example.com/', { waitUntil: 'networkidle0', timeout: 60_000 });
  await page.screenshot({ path: 'example-full.png', fullPage: true });
} finally {
  await browser.close();
}

For an element, use a selector and screenshot the corresponding element handle or locator following the current Puppeteer screenshot guide. Its guide notes that Puppeteer attempts to scroll a hidden element into view. As with Playwright, schedule the script separately and use an explicit, bounded readiness condition appropriate to the page.

5. Schedule locally or in CI

For a machine you control, cron can start a capture command at an interval. For a repository-based workflow, GitHub Actions provides a scheduled event and a convenient place to keep the workflow beside the capture configuration. Other schedulers can work too; choose based on where the browser runs, how you store outputs, and how you observe failures.

  • Use a schedule interval appropriate to the value of each image and the cost of browser runtime.
  • Make the capture job repeatable: pin or otherwise control browser and dependency versions where your setup allows it.
  • Write outputs to a known location and define retention or review rules.
  • Provide a manual trigger for debugging and an easy way to rerun failed captures.
  • Record run time and result status so missing images are distinguishable from unchanged pages.

GitHub Actions is not a real-time scheduler. Scheduled runs execute on the default branch and may be delayed or dropped during high load. If every run must happen within a strict time window, use a scheduler whose documented guarantees match that requirement, and add monitoring for missed work.

6. Make screenshots repeatable

For a simple archive, repeatability means using the same capture settings and consistently naming outputs. For visual regression, it also means controlling the browser, operating system, viewport, fonts, locale, and other rendering inputs. Playwright warns that rendering can vary with host OS, browser version, settings, hardware, power source, headless mode, and other factors; its visual comparison guidance recommends using the same environment that generated the baselines.

  • Keep viewport dimensions fixed; responsive layouts can change at breakpoints.
  • Wait for a meaningful page condition instead of relying on arbitrary short sleeps.
  • For dynamic content, decide whether to mask, stabilize, or accept changing regions before comparing images.
  • Use a stable runner image and browser setup for baselines and scheduled runs.
  • Review changed images in context; a pixel difference alone does not explain whether the page change is meaningful.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. For a one-off capture, call its API with a URL; see the API documentation. This cURL request saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. The API supports caching, async jobs, and bulk capture for up to 100 URLs per call, which can help when building a recurring capture workflow.

Sign up free for 1,000 screenshots a month with no card.

8. Troubleshooting

Symptom Likely cause What to do
Workflow does not start at the expected minute Scheduled CI runs can be delayed under load; cron is not precise timing. Check the workflow run history, allow for delay, and use a scheduler with suitable guarantees if timing is strict.
No scheduled run appears The workflow may not be on the default branch, or a run may have been delayed or dropped. Confirm the workflow is present on the default branch and inspect the scheduler’s status and event logs.
Browser executable is missing The runner installed the automation package without its required browser. Install the browser dependency as part of setup; shot-scraper’s documented workflow includes installing its browser.
Capture times out The page is slow, waiting for all network activity never completes, or the configured timeout is too short. Use a bounded timeout and wait for a relevant selector or page state when network-idle is inappropriate.
Screenshot is blank or incomplete The capture happened before the page rendered or before lazy content entered the viewport. Wait for a page-specific readiness condition; for full-page captures, confirm the tool and page behavior load the content you need.
Diffs appear without a meaningful page change Rendering inputs or dynamic content changed between baseline and run. Use the same runner environment, browser version, viewport, and stable page state; review dynamic regions.
Workflow captures images but does not save them Output paths or repository write permissions are wrong, or the commit step has no changes to commit. Check the capture output path and job permissions. A no-change run is expected to create no commit.
Repository grows quickly Every run adds or updates large binary files without a retention plan. Choose an output store and retention policy suited to the archive; avoid keeping every image in source control unless that is intentional.

9. Performance, reliability, and cost

Browser captures consume time and resources in proportion to the pages, browser work, and frequency of runs; the reviewed sources provide no controlled benchmark for comparing these tools, so test your own URLs and runner configuration before setting an interval. A one-page archive and a large multi-URL capture have different operational needs.

  • Performance: limit parallel browser work to what the runner can handle, avoid unnecessary recaptures, and set timeouts so a stuck page does not hold a job forever.
  • Reliability: distinguish a failed capture from an unchanged page, retain logs, and provide retries or reruns where appropriate. Account for scheduler delays and missed runs.
  • Cost: self-managed tools have no per-shot API price in the reviewed sources, but scheduled runners, compute time, storage, and maintenance still have costs. Estimate from your own run frequency and retention needs.
  • Visual checks: keep baseline and capture environments consistent; otherwise a diff may reflect rendering variation rather than a site change.

10. Frequently asked questions

Can GitHub Actions take screenshots on a schedule without a separate server?

Yes. A workflow can install the browser tooling, run captures, and store results on a hosted runner. GitHub documents scheduling caveats, so it should not be treated as a precision timer.

Which tool should I use for a simple list of pages?

Start with shot-scraper if a CLI and configured captures meet your needs. Its documentation includes a multi-capture configuration and a scheduled workflow pattern.

Which is best for visual regression tests?

Playwright has documented screenshot assertion support. Keep the baseline and test runner environment consistent, then review diffs that matter to the application.

Is Puppeteer only for full-page images?

No. Its screenshot guide covers selected-element screenshots as well as page screenshots.

Can a scheduled screenshot archive guarantee that every run happens on time?

Not with GitHub Actions schedules: its documentation allows for delays and dropped queued runs under high load. Select a scheduler based on the timing guarantees your use case requires.