How to Stagger Scheduled Screenshots Across a List of Websites
Schedule website screenshots as separate, staggered jobs. Build a Playwright worker with per-site readiness checks, traceable outputs, retries, and controlled concurrency.
To stagger scheduled screenshots across a list of websites, give each URL its own capture job and planned start time, then have a scheduler release those jobs gradually. The browser script should handle navigation, a page-specific readiness condition, the screenshot, and a status record. The scheduler controls when work starts; it does not need to know how a browser captures a page.
This separation avoids launching every browser at once, makes failures attributable to individual sites, and lets you set concurrency and retry policies. There is no universal safe interval: choose a cadence that fits your workload, scheduler behavior, and each site’s terms and rate limits.
1. Choose the capture scope and readiness condition
Before writing the schedule, decide what a successful capture means for each site. A screenshot can represent the visible viewport, the full page, or a particular element. Use the narrowest scope that answers your monitoring question: viewport for the initial visitor view, full page for below-the-fold changes, or an element for a specific component. Playwright documents these screenshot forms in its screenshot guide.
Readiness should also be site-specific. Waiting for a meaningful selector or application state is generally more useful than applying the same fixed sleep to every URL. For example, a product page may be ready when its price element appears, while a dashboard might need a chart container. A selector is only an example; verify it against the site you capture.
2. Model each URL as an independent job
Store a stable identifier, URL, readiness selector, and optional schedule offset for every target. Keep the offset separate from the capture code so you can adjust cadence without changing browser behavior.
[
{
"id": "docs-home",
"url": "https://example.com/docs",
"readySelector": "main",
"offsetSeconds": 0
},
{
"id": "store-home",
"url": "https://shop.example.org/",
"readySelector": "h1",
"offsetSeconds": 60
}
]
The offsets above illustrate configuration only; they are not a recommended universal interval. In production, use a database or scheduler queue if operators need to edit targets without changing code. Validate that IDs are unique, URLs use expected schemes, and selectors are present where required.
3. Build a Playwright capture worker
This runnable Node.js example reads a JSON target list, processes targets sequentially, waits for each configured selector, takes a full-page screenshot, and writes a JSON status record even when an individual target fails. Sequential execution is the simplest way to prevent a burst. If you later add concurrency, cap it deliberately and consider a per-host limit.
npm init -y
npm install playwright
Save the target list as targets.json, then save this as capture.mjs:
import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
const targets = JSON.parse(await readFile(new URL('./targets.json', import.meta.url), 'utf8'));
const outputDir = new URL('./captures/', import.meta.url);
await mkdir(outputDir, { recursive: true });
function safeId(value) {
return value.replace(/[^a-zA-Z0-9_-]/g, '_');
}
function stamp(date = new Date()) {
return date.toISOString().replaceAll(':', '-').replaceAll('.', '-');
}
const browser = await chromium.launch({ headless: true });
const runStamp = stamp();
const results = [];
try {
for (const target of targets) {
const startedAt = new Date().toISOString();
const base = `${safeId(target.id)}-${runStamp}`;
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
const response = await page.goto(target.url, {
waitUntil: 'domcontentloaded',
timeout: 45000
});
if (target.readySelector) {
await page.locator(target.readySelector).waitFor({ state: 'visible', timeout: 20000 });
}
await page.screenshot({ path: new URL(`${base}.png`, outputDir).pathname, fullPage: true });
results.push({
id: target.id,
url: target.url,
status: 'success',
httpStatus: response?.status() ?? null,
startedAt,
finishedAt: new Date().toISOString(),
image: `${base}.png`
});
} catch (error) {
results.push({
id: target.id,
url: target.url,
status: 'failure',
startedAt,
finishedAt: new Date().toISOString(),
error: String(error)
});
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
await writeFile(
new URL(`run-${runStamp}.json`, outputDir),
JSON.stringify(results, null, 2)
);
console.log(JSON.stringify(results, null, 2));
Run it with node capture.mjs. Install the browser binary in the environment that runs the worker if needed; Playwright’s setup instructions cover browser installation. This example intentionally uses domcontentloaded followed by a page-specific selector. Use a different readiness condition if the target renders meaningful content later.
The code runs the list once. The scheduler is responsible for invoking it repeatedly or enqueueing individual target jobs at their due times. If you need actual spacing within a run, schedule each target as a separate job or add an explicit delay between jobs. A single run with a sequential loop avoids simultaneous starts but does not guarantee exact start times.
4. Schedule jobs at staggered times
There are two common scheduling patterns:
- Per-URL scheduled jobs: create one recurring job per target, each with its own offset. This gives each site an independent schedule and failure history.
- Queue and paced dispatch: add all due URLs to a queue, then have a worker release them at a controlled rate. This is useful when the list changes often or the scheduler has job-count limits.
For a small fixed list, a scheduler can start one worker at a time, with each worker selecting the next due URL. For a larger or dynamic list, persist next_due_at per URL, claim due rows atomically, and enqueue one capture task per row. A typical record includes URL ID, due time in UTC, status, attempt count, last error, and output key.
Before depending on any scheduler, confirm its official behavior for timezone interpretation, delayed starts, missed runs, concurrent executions, retries, and maximum runtime. Exact start-time precision and missed-run behavior depend on the chosen scheduler; do not assume a cron expression guarantees a particular execution second.
5. Make output names and records traceable
Use a predictable storage key containing the target ID and capture timestamp, such as captures/docs-home-2026-10-04T12-30-00-000Z.png. Store the original URL, start and finish times, status, HTTP response status when available, and error details beside the image. This distinguishes a missing capture from a page that loaded successfully but looked unchanged.
For long-term monitoring, write the image and status record to durable storage rather than relying only on a worker’s temporary filesystem. Define retention and cleanup separately from capture scheduling, and avoid placing secrets in filenames or logs.
6. Keep visual comparisons stable
Visual results can change because of differences in host OS, browser version, browser settings, hardware, power source, or headless mode. Playwright documents these sources of rendering variation in its visual comparisons guidance. Keep the browser version, operating environment, viewport, device scale factor, locale, and color scheme consistent where comparisons matter.
Also account for page content that changes naturally, such as timestamps, rotating promotions, ads, animations, and personalized content. Playwright screenshot options support style overrides; a narrowly scoped stylesheet can hide known volatile elements when doing so does not conceal the content you intend to monitor. Use the same override on every run and record it with the capture configuration.
7. Handle retries, rate limits, and partial failures
- Record failure per URL and continue processing the rest of the list.
- Retry transient navigation failures with a bounded attempt count and backoff. Avoid rapid retry loops that multiply load on a struggling site.
- Set a per-host concurrency ceiling. A global worker limit alone can still send many simultaneous requests to one domain.
- Keep the original due time and actual start time so delayed jobs are visible in reports.
- Make output writes idempotent or use attempt-specific keys. A retry should not silently overwrite a successful result unless that is the intended retention policy.
- Respect robots policies, terms, authentication requirements, and rate limits that apply to the sites you capture.
There is no source-backed interval that is safe for every site. Base spacing and concurrency on the workload, site requirements, scheduler capabilities, and observed failures.
8. Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Every screenshot starts together | The scheduler launches one task for the whole list, or the worker fans out all URLs concurrently. | Use per-URL due times, a paced queue, or a sequential worker. Add an explicit concurrency cap. |
| Screenshot is blank or incomplete | The page had not rendered its useful content when capture began, or the readiness selector is wrong. | Wait for a meaningful visible selector or application state. Check the selector in the page and increase its timeout only when the page legitimately needs longer. |
| Navigation times out | The site is slow, unreachable, waiting on long-lived requests, or the timeout is too short. | Use a navigation condition suited to the page, inspect the failing URL, and set a bounded timeout. Do not treat a larger timeout as a cure for a permanently unavailable site. |
| Some URLs are skipped after one error | An exception escapes the loop or the scheduler treats the whole batch as one indivisible task. | Catch errors per target, persist each result, and let the remaining jobs proceed. |
| Duplicate captures appear | A scheduled run overlaps a previous run, or retries create duplicate tasks. | Use a per-target lock or idempotency key and define whether overlapping due jobs should queue, coalesce, or be skipped. |
| Visual diffs fluctuate between runs | Browser or host environment changed, or page content is dynamic. | Pin the rendering environment and viewport; use stable readiness checks and consistent style overrides for known volatile regions. |
| Scheduled time differs from expectation | Timezone, daylight saving, delayed execution, or missed-run behavior was misunderstood. | Store intended due times in UTC where practical and verify the selected scheduler’s timezone and catch-up semantics. |
9. Performance, reliability, and cost
Launching a browser has overhead, so a shared worker that processes a few jobs sequentially may use fewer resources than starting a fresh browser process for every URL. Reusing a browser while creating an isolated page or context per target can reduce startup work, but isolate cookies and storage when sites or accounts must not share state. For larger workloads, measure queue delay, browser memory, capture duration, and failure rate before raising concurrency.
Staggering smooths resource demand; it does not make total work disappear. Estimate total run time from the number of URLs and observed per-site capture time, then leave enough slack for slow pages and bounded retries. Keep scheduler runtime limits in view. A queue with independently acknowledged jobs is more resilient to worker interruption than a single long batch.
Self-managed cost includes compute, storage, scheduler usage, and engineering time for browser updates, logs, retries, and retention. A managed screenshot service may combine multi-URL capture and scheduling, but confirm its current scheduling features, limits, storage options, and pricing before choosing it. The research sources do not establish a feature-by-feature comparison or prices for other services.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-call API can capture a URL while your scheduler controls when each request is sent. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Sign up for 1,000 free screenshots a month with no card.
FAQ
Should I use a fixed delay between every URL?
Only if a fixed cadence fits the sites and workload. Per-URL due times are better when targets have different schedules or host-specific constraints; a queue is useful when the list changes frequently.
Does staggering make screenshots visually consistent?
No. Staggering controls when captures run. Consistency also depends on keeping browser and rendering settings stable and accounting for dynamic page content.
Should the scheduler retry an entire batch?
Usually, track and retry individual URL jobs so a failure on one site does not repeat successful captures for the whole list.
Can I use full-page capture for every site?
Yes, when below-the-fold content is part of what you monitor. For a specific component or initial viewport, a narrower capture is smaller and easier to interpret.


