Best Open-Source Tools for Scheduling Website Screenshots on a Server
Compare shot-scraper, Playwright, and self-hosted schedulers, then set up repeatable website screenshots on a server with runnable examples.
For a compact open-source setup, use shot-scraper to capture pages and run it on a schedule. It supports command-line and multi-page YAML captures, and its documentation includes a GitHub Actions workflow. Choose Playwright when each page needs custom browser actions or application-specific waiting. A scheduler such as cron, a workflow runner, or a self-hosted job manager determines when either tool runs; the capture tool and scheduler are separate parts of the system.
If you would rather call a screenshot service than install and operate a browser, ScreenshotNeo is the first alternative to try: it removes common consent banners, popups, and chat widgets before capture, and only clean shots are billed.
1. Pick the capture and scheduling layers
| Option | Capture | Scheduling | Good fit | Tradeoff |
|---|---|---|---|---|
| shot-scraper | CLI and YAML multi-shot configuration; full-page and element captures | GitHub Actions, cron, or another job runner | Repeated URL lists with straightforward settings | You choose the scheduler and output storage |
| Playwright | Browser automation API for viewport, full-page, and element screenshots | Run your script from cron, a workflow, or another scheduler | Custom interactions, waits, or an existing Playwright project | You write and maintain more capture code |
| Cronbase plus a capture tool | Use shot-scraper or your Playwright script | Self-hosted cron dashboard | Operators who want a UI for defining and monitoring jobs | Cronbase is a general-purpose scheduler; browser setup remains yours |
Cronbase describes itself as an open-source, self-hosted cron manager. The reviewed project information does not establish a screenshot-specific integration: treat it as the scheduler and verify that its job environment can run your chosen browser tool.
2. Set up shot-scraper for a URL list
shot-scraper is a good default when the job is mostly “capture these pages and save the files.” Its stable documentation describes installing the Python package and then installing its browser.
- Install Python and create an isolated environment on the machine or runner that will execute the job.
- Install shot-scraper and its browser.
- Make a small configuration with URLs and output names.
- Run it manually, inspect the images, and only then add the command to a scheduler.
python -m venv .venv
. .venv/bin/activate
python -m pip install shot-scraper
shot-scraper install
For Windows PowerShell, activate with .venv\Scripts\Activate.ps1. Keep the environment and browser installation available to the scheduled job; a cron process may not use the same PATH or working directory as your interactive shell.
Create shots.yml using shot-scraper’s documented multi-shot configuration format, for example:
- url: https://example.com/
output: screenshots/example-home.png
- url: https://example.org/
output: screenshots/example-org.png
Run the configuration with:
mkdir -p screenshots
shot-scraper multi shots.yml
Check the installed version’s stable documentation for the current YAML fields and CLI options before adding advanced settings. The command-line tool also supports individual captures; use its screenshot reference for available viewport and element capture options.
3. Schedule the command
Option A: cron on a Linux server
Put the capture invocation in a small script so you can set its working directory, log output, and return status explicitly. Save as /opt/site-shots/run.sh and adjust the paths:
#!/usr/bin/env bash
set -euo pipefail
cd /opt/site-shots
. .venv/bin/activate
mkdir -p screenshots logs
shot-scraper multi shots.yml
Make it executable with chmod +x /opt/site-shots/run.sh, run it manually once, then add a cron entry. This example runs every day at 06:15 in the server’s configured timezone:
15 6 * * * /opt/site-shots/run.sh >> /opt/site-shots/logs/cron.log 2>&1
Confirm the server timezone and cron service behavior for your operating system. If exact UTC timing matters, configure the host accordingly and document the intended timezone. Prevent overlapping runs if one capture cycle can take longer than the interval; a lock mechanism such as flock can be added where available.
Option B: GitHub Actions
shot-scraper documents a workflow pattern that installs the package and browser, runs configured captures, and commits the results to the repository. A workflow schedule determines cadence; the example in the project documentation is a pattern, not a guarantee of a particular schedule or start time. Scheduled workflow runs are appropriate when repository history is a useful output destination; use a server scheduler or separate storage when it is not.
Follow the current shot-scraper GitHub Actions guide for its workflow structure. Set the workflow schedule explicitly, retain only the files you need, and check the workflow logs for failures. Schedule syntax and runner behavior belong to GitHub Actions and can change; verify against its current documentation when deploying.
Option C: a self-hosted job dashboard
With Cronbase or another self-hosted scheduler, configure the scheduled command to invoke your capture script. Confirm the job user, working directory, environment variables, browser dependencies, output permissions, and logs. A dashboard can manage general commands, but your paired capture tool still handles page navigation and screenshots.
4. Use Playwright when pages need browser actions
Playwright is a better fit when a page must be navigated, interacted with, or waited on in a way that does not fit a simple capture configuration. The following Node.js example creates a project, installs Chromium, and captures one viewport screenshot per URL.
mkdir scheduled-shots
cd scheduled-shots
npm init -y
npm install playwright
npx playwright install chromium
Save as capture.mjs:
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
const targets = [
{ url: 'https://example.com/', file: 'screenshots/example-home.png' },
{ url: 'https://example.org/', file: 'screenshots/example-org.png' },
];
await mkdir('screenshots', { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
const context = await browser.newContext({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
});
const page = await context.newPage();
for (const target of targets) {
await page.goto(target.url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
// Prefer a meaningful page condition when the content is application-rendered.
await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 }).catch(() => {});
await page.screenshot({ path: target.file, fullPage: true, animations: 'disabled' });
}
} finally {
await browser.close();
}
Run it with node capture.mjs. Replace main with a selector that reliably indicates readiness for your pages. The sample tolerates a missing main selector, but you should remove that fallback or use a more specific readiness condition if capturing too early would produce a misleading image.
The Playwright screenshot API supports viewport screenshots by default, fullPage: true for the full scrollable page, and element screenshots through a locator such as await page.locator('.header').screenshot({ path: 'header.png' }). Other relevant options include output path or returned image bytes, image type, quality for supported formats, animation handling, and transparent background where supported. Consult the Page screenshot API for current option details and constraints.
To capture one element, use a locator after the page is ready:
const card = page.locator('[data-testid="pricing-card"]');
await card.waitFor({ state: 'visible' });
await card.screenshot({ path: 'screenshots/pricing-card.png' });
To run this script periodically, use the same cron or job-runner pattern as above, changing the invoked command to node /opt/site-shots/capture.mjs. Install the Playwright browser in the job environment and run the script as the same user that owns the output directory.
5. Make scheduled captures comparable and dependable
- Keep the runtime consistent. Pin the application dependencies and control the browser build, operating system, fonts, viewport, device scale factor, locale, and color scheme as far as your deployment allows. Playwright warns that screenshot rendering can vary with host OS, browser version, settings, hardware, and headless mode. See its visual comparisons guidance.
- Wait for page state, not an arbitrary long delay. Use a visible selector or a page-specific ready signal when content loads asynchronously. A fixed delay can be too short on a slow run and waste time on a fast one.
- Choose viewport or full page deliberately. Full-page images can be very tall and may expose sticky-header or lazy-loading behavior. For a single component, capture a locator instead.
- Make outputs traceable. Use stable names for latest-state snapshots or include a date in the name when keeping history. Select a retention policy and storage location appropriate to the images.
- Keep secrets out of source control. If a target requires authentication, inject credentials through the scheduler’s secret mechanism and avoid logging them. Ensure you are authorized to capture the target and store its contents.
- Record useful logs. Log the target, start/end time, exit status, and error message. Alert on repeated failures if these images support operational decisions.
- Control concurrency. Start with sequential captures. Add parallel workers only when server resources and target sites can handle the extra browser load; otherwise runs can compete for memory or overload the sites.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Command works in a terminal but cron cannot find it | Cron has a limited environment or different working directory | Use absolute paths, explicitly activate the virtual environment, and set the working directory in the script. |
| Browser executable missing | Browser was not installed in the scheduled user’s environment, or the runtime changed | Run the tool’s browser installation step as the job user and keep installation aligned with the installed tool version. |
| Permission denied writing images or logs | Job user does not own the destination directory | Create the directory and set ownership/permissions for the scheduler’s user; avoid writing into a developer’s home directory by assumption. |
| Screenshot is blank or has a loading spinner | Capture happened before client-rendered content was ready, navigation failed, or the page rejected automated traffic | Check logs and response state; wait for a page-specific selector; verify the URL is reachable from the server. Do not treat a bot check as a valid page capture. |
| Images differ between runs | Dynamic content, animation, changing data, fonts, browser, OS, or viewport changed | Stabilize the environment and page state, disable animations where suitable, and avoid comparing captures from different runtime configurations. |
| Job overlaps or runs out of memory | A capture cycle takes longer than the schedule interval or launches too many browsers | Use a lock or concurrency limit, capture sequentially, reduce the batch, or lengthen the interval. |
| Full-page output is unexpectedly large or incomplete | The page is unusually long, lazy-loaded content is not ready, or the output target is unsuitable | Wait for relevant content, consider element or viewport capture, and check the image dimensions and storage limits. |
| Scheduled workflow does not run at the expected minute | Workflow schedules are controlled by the workflow platform and may not start exactly when configured | Check the workflow platform’s current scheduling guidance and use a server scheduler if the required cadence or control differs. |
7. Performance, reliability, and cost
For self-hosted captures, the main costs are the server or runner time, browser storage and memory, output storage, and the maintenance work of keeping the runtime usable. The dossier provides no comparative benchmarks or universal runtime figures, so estimate from your own pages and environment. Start with a small batch, measure the actual job duration and peak resource use, then set a schedule with room for the slowest expected run.
Reliability depends on the whole job path: scheduler invocation, browser availability, target reachability, page readiness, write permissions, and output retention. Return a failing process status when a required capture fails, preserve logs, and decide whether one failed URL should stop the batch or allow later URLs to continue. For visual comparisons, keep browser and host conditions stable; exact pixel identity across different environments is not guaranteed.
GitHub Actions can store image changes in repository history as in the documented shot-scraper workflow. This provides version history but grows the repository as images accumulate; choose retention and storage based on how long you need historical captures.
Or skip the browser setup
ScreenshotNeo provides a hosted screenshot API and MCP server. A single GET request captures a URL as an image or PDF; its docs list additional capture controls and integrations. The call below saves a WebP capture of Stripe:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. The practical reasons to consider it are specific: cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, while paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.
FAQ
Does shot-scraper schedule screenshots by itself?
It performs captures. You choose a scheduler such as GitHub Actions, cron, or a job manager and configure when it invokes the capture command.
Should I use cron or GitHub Actions?
Use cron when you control an always-on server and want jobs and files there. Use a hosted workflow when repository-based execution and image history fit your needs. The right choice depends on where the job should run and where its output belongs.
Can Playwright take full-page and element screenshots?
Yes. Its screenshot API supports full-page capture and locator-based element capture. The official screenshot guide has examples.
Will the same page always produce an identical image?
No. Page content and rendering environments can vary. Control the browser, host, viewport, and page state when consistency matters.
Do I need a dedicated physical server?
No specific hardware is established as necessary by the reviewed sources. Run the job on an existing server or workflow runner that can install and execute the chosen browser tool.


