How to Retry Scheduled Website Screenshots When a Page Times Out
Retry timed-out screenshot jobs safely with bounded attempts, useful logs, and idempotent output. Includes a runnable Python Playwright worker and a ScreenshotNeo option.
When a scheduled screenshot times out, mark that capture attempt as failed and let the scheduler or job wrapper make a finite number of fresh attempts. Record the URL, run ID, attempt, elapsed time, timeout stage, and error. Save or publish an image only after a successful capture, using a stable output name so retries do not create confusing duplicates.
Playwright Test retries, screenshot assertion retries, and retries for a recurring capture job are separate mechanisms. If your scheduled task is a standalone script, implement the retry policy in that script or its scheduler; Playwright Test’s retries setting does not configure cron or another scheduler. The retry count and delay depend on your schedule and site; there is no universal formula.
1. Choose the right retry layer
| What failed | What retries it | Use it when |
|---|---|---|
| Page navigation or readiness timed out | Your capture worker or scheduler | A recurring script or API job needs a fresh capture attempt. |
| A Playwright Test failed | Playwright Test’s retries configuration |
The scheduled command runs a Playwright Test suite and retrying the test is acceptable. |
| A visual snapshot assertion did not stabilize or match | expect(page).toHaveScreenshot()‘s assertion timeout |
You are testing a visual baseline inside Playwright Test, rather than retrying a recurring capture job. |
Playwright Test’s per-test timeout defaults to 30 seconds, and its failed-test retries option is separate from scheduler behavior. Playwright Test timeout documentation and configuration reference describe these settings. A screenshot assertion waits for two consecutive page screenshots to match before comparing the final image; its timeout controls assertion retrying, and the assertion is available only in the Playwright test runner. It is not a recurring-job retry policy. Playwright screenshot assertion documentation
2. A reliable retry sequence
- Set an explicit navigation timeout and a separate readiness condition. Avoid waiting indefinitely: one slow page can consume the whole schedule window.
- Classify the failure. Navigation timeout, readiness selector timeout, browser launch error, invalid URL, access denied, and file write error need different responses.
- Retry transient errors only, with a small finite attempt limit and increasing delay. Choose limits that fit the interval before the next scheduled run.
- Use a new page for each attempt. Close it even after failure so resources do not accumulate.
- Write to a temporary file and atomically replace the destination only after capture succeeds. Preserve an earlier good image when a later attempt fails.
- Prevent overlapping runs for the same URL set, or use run-specific temporary paths and a lock.
- Log every attempt and alert or mark the job failed after the final attempt. A missing image must not appear as a successful result.
3. Runnable Python Playwright worker
This example uses exponential backoff with a finite attempt limit. It captures one URL per invocation, saves to a temporary file, and replaces the stable output only after a successful screenshot. Install Playwright with pip install playwright and install its browser with playwright install chromium. Set TARGET_URL and run with Python 3. The scheduler can invoke the script at its normal cadence.
import asyncio
import json
import os
import random
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright
URL = os.environ.get("TARGET_URL", "https://example.com")
OUTPUT = Path(os.environ.get("OUTPUT", "site.png"))
RUN_ID = os.environ.get("RUN_ID", datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ"))
MAX_ATTEMPTS = int(os.environ.get("MAX_ATTEMPTS", "3"))
NAVIGATION_TIMEOUT_MS = int(os.environ.get("NAVIGATION_TIMEOUT_MS", "20000"))
READY_SELECTOR = os.environ.get("READY_SELECTOR", "")
def log(**fields):
print(json.dumps({"time": datetime.now(timezone.utc).isoformat(), **fields}), flush=True)
def validate_url(url):
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
raise ValueError(f"Invalid HTTP(S) URL: {url!r}")
async def capture_once():
started = time.monotonic()
stage = "browser_start"
temporary = OUTPUT.with_name(OUTPUT.name + f".{RUN_ID}.tmp")
try:
async with async_playwright() as p:
browser = await p.chromium.launch()
try:
page = await browser.new_page(viewport={"width": 1440, "height": 1000})
page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
stage = "navigation"
response = await page.goto(URL, wait_until="domcontentloaded")
if response and response.status >= 400:
raise RuntimeError(f"HTTP status {response.status}")
if READY_SELECTOR:
stage = "readiness"
await page.locator(READY_SELECTOR).wait_for(state="visible", timeout=10000)
stage = "screenshot"
await page.screenshot(path=str(temporary), full_page=True, timeout=15000)
finally:
await browser.close()
stage = "publish"
temporary.replace(OUTPUT)
return {"ok": True, "path": str(OUTPUT), "elapsed_ms": round((time.monotonic() - started) * 1000)}
except Exception:
temporary.unlink(missing_ok=True)
raise
async def main():
validate_url(URL)
if MAX_ATTEMPTS < 1:
raise ValueError("MAX_ATTEMPTS must be at least 1")
for attempt in range(1, MAX_ATTEMPTS + 1):
started = time.monotonic()
try:
result = await capture_once()
log(run_id=RUN_ID, attempt=attempt, url=URL, outcome="success", **result)
return
except Exception as error:
elapsed_ms = round((time.monotonic() - started) * 1000)
log(run_id=RUN_ID, attempt=attempt, url=URL, outcome="failure",
elapsed_ms=elapsed_ms, error_type=type(error).__name__, error=str(error))
if attempt == MAX_ATTEMPTS:
raise SystemExit(f"Capture failed after {attempt} attempts: {error}")
delay = min(30, 2 ** (attempt - 1)) + random.uniform(0, 0.5)
await asyncio.sleep(delay)
if __name__ == "__main__":
asyncio.run(main())
In this worker, the example defaults are implementation choices, not recommended values for every site. Tune timeout, attempt limit, backoff, readiness selector, and output naming to the job’s cadence and freshness requirement. The exception handler retries broadly for illustration; production code should classify failures and avoid retrying permanent errors such as malformed URLs or known access-denied responses.
4. Timeouts, readiness, and Playwright Test settings
Navigation completion is not visual readiness
page.goto() can wait for commit, domcontentloaded, load, or networkidle. The default is load; Playwright discourages using networkidle as a general testing readiness signal. For a screenshot, use a meaningful page-specific condition such as a visible content selector when possible. Navigation timeout and selector timeout are distinct stages. See the Page API.
If the scheduled command runs Playwright Test
Set test timeout and retries in the test config. These retry failed tests, so the whole test may run again; they do not retry only a URL within an independent scheduler job.
import { defineConfig } from '@playwright/test';
export default defineConfig({
timeout: 60_000,
retries: 2,
reporter: [['list'], ['html', { open: 'never' }]],
});
Use a test timeout that leaves room for setup, navigation, readiness, capture, and cleanup. Avoid stacking a long test timeout with long internal retries if the scheduler has a shorter deadline.
5. Scheduling, idempotency, and overlapping runs
- Stable identity: assign a run ID and a stable destination per page or capture slot. Decide whether a retry belongs to the same run or creates a new scheduled slot.
- Atomic publication: capture to a temporary file on the same filesystem, then rename after success. If artifacts go to object storage, upload to a temporary key and update a manifest/pointer only after upload completes.
- Concurrency: use a scheduler concurrency limit, distributed lock, or per-URL lease so two runs do not overwrite each other. Expire locks safely if workers crash.
- Freshness: if a retry succeeds after the next scheduled run begins, define which run is authoritative. Include scheduled time in metadata and avoid silently replacing a newer artifact with an older run.
- Exhaustion: exit nonzero or emit the scheduler’s failure signal after attempts are exhausted. Alert on final failures rather than every transient attempt unless operators need immediate visibility.
6. cURL and JavaScript examples for scheduled API captures
If a hosted screenshot API performs the browser capture, place retries around the HTTP request in the scheduled worker. A timeout from the API request is ambiguous: the server may have completed the work even though the client missed the response. Use a stable run identifier in your own job records, and avoid publishing an artifact until you have a successful response. Follow the API’s documented semantics for any request identifier or idempotency feature.
cURL attempt
curl --fail --show-error --silent --max-time 90 \
-G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o screenshot.tmp
# Move screenshot.tmp to the stable destination only when curl succeeds.
mv screenshot.tmp screenshot.webp
Have the scheduler rerun the command on a transient nonzero exit, subject to its finite retry policy. Configure its execution timeout to exceed the request timeout enough to allow cleanup.
Node.js request attempt
const q = new URLSearchParams({
access_key: process.env.SCREENSHOTNEO_API_KEY,
url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`, {
signal: AbortSignal.timeout(90_000)
});
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await fs.promises.writeFile('screenshot.tmp', bytes);
await fs.promises.rename('screenshot.tmp', 'screenshot.webp');
Import fs with import fs from 'node:fs'; in an ES module. Keep API keys in environment variables or a secrets manager, not in source control or logged URLs.
7. Troubleshooting
| Symptom | Likely cause | Action |
|---|---|---|
page.goto timeout |
Slow server, network trouble, or waiting for a later load event that never occurs. | Log the stage and elapsed time. Choose an appropriate navigation condition and explicit timeout; retry transient failures with a fresh page. |
| Navigation succeeds but screenshot is blank or incomplete | The page rendered a shell before client-side content, or a required resource is delayed. | Wait for a page-specific selector or readiness signal before capture. Treat empty output as a failed attempt if blank pages are not acceptable. |
| Readiness selector times out | Selector changed, content is conditional, or the page did not reach the expected state. | Check the selector and page state. Keep readiness timeout separate from navigation timeout. |
| Retry never happens | The retry is configured in Playwright Test but the job is a plain script, or the scheduler does not retry nonzero exits. | Put retry logic in the process that actually launches the capture and verify how that scheduler handles failures. |
| Duplicate or stale images | Overlapping runs, retry-specific filenames, or late completion overwriting newer output. | Use run IDs, a lock or concurrency limit, atomic replacement, and a freshness rule. |
| Partial file replaces the last good image | Writing directly to the published destination before capture completes. | Write to a temporary path and publish only on success. |
| Worker memory or browser processes accumulate | Pages or browsers are not closed after an exception, or retries happen inside one long-lived page. | Close browser resources in a finally block and create a new context/page for each attempt. |
| Every retry fails on a 4xx or access-denied page | The failure is permanent or requires authentication/allowlisting. | Classify it and stop retrying blindly; correct the URL or access configuration. |
8. Performance, reliability, and cost
Retries consume browser time, network capacity, and scheduler slots. Bound both attempts and per-attempt timeouts so a slow URL cannot starve later work. Backoff with jitter helps avoid synchronized retries when many URLs fail together. Capture only the pages and content needed, reuse browser processes only when you can isolate page state safely, and close each page/context. For large URL sets, limit concurrency to match available CPU, memory, and the target site’s tolerance.
Retain enough logs to distinguish slow navigation from readiness failures, browser startup issues, and storage problems. Record run ID, attempt number, scheduled time, URL, timeout stage, elapsed time, result, and artifact path. Do not claim a retry succeeded merely because the scheduler relaunched the job; verify a successful capture and publish step.
With a browser you operate, retries add compute and network usage. With a hosted service, billing and retry semantics vary by provider; check that provider’s documentation rather than assuming failed requests are free or safe to replay.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. One GET request returns a PNG, JPEG, WebP, or PDF; see the API documentation for options and response details. For scheduled jobs, put the request in your scheduler’s bounded retry wrapper and publish only successful responses.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
- Cookie and consent banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
- An MCP server gives AI agents tools to take screenshots, inspect page info, and capture PDFs.
- The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card.
FAQ
Should I retry a screenshot timeout immediately?
Usually, wait briefly and use a finite backoff. An immediate retry can hit the same temporary outage and synchronize with other failing jobs.
Does toHaveScreenshot() retry a scheduled capture?
No. It retries visual stabilization and comparison within Playwright Test. A recurring capture needs retry orchestration around the scheduled worker or API call.
Should a retry overwrite the last good screenshot?
Only after a new capture completes successfully and your freshness policy says it is the authoritative result.
What should happen after the final attempt?
Mark the run failed, preserve the last known good artifact, and make the failure visible through logs or an alert.


