ScreenshotNeo

BlogHow-to

How to Retry Failed URLs in a Batch Screenshot Job

Retry only the URLs that failed, preserve completed captures, and keep attempts bounded. This guide includes a runnable Python batch workflow and recovery checklist.

By the ScreenshotNeo team4 October 202611 min read

Retry failed URLs individually and checkpoint every successful capture. A batch is a collection of separate work items: if one page times out, keep the completed screenshots and retry only that page, with a maximum attempt count and a recorded reason for each failure. Retrying the entire job is safe only when completed work is persisted and the job can skip it.

This guide uses Python, Pillow, and Playwright to demonstrate a local batch workflow. The retry policy is deliberately conservative: it retries timeouts and selected network failures, but leaves invalid URLs and other permanent failures for review. Your screenshot provider may expose different errors, so adapt classification to its documented behavior.

1. Choose the retry unit

First decide what failed. A URL capture, a test that checks a screenshot, and a scheduler job are different units. Application code can retry one URL. A scheduler retry may rerun the whole job, including URLs that already succeeded, unless the batch stores per-URL completion state.

Failure scope What to retry Key safeguard
One navigation or capture That URL, if the error is transient Bound attempts and preserve the last error
A visual snapshot mismatch The screenshot assertion according to the test framework Keep the browser environment stable
A scheduler job failure The job only if URL-level progress is checkpointed Skip already completed captures on restart

Playwright’s visual screenshot assertion behavior is separate from URL recovery: its documentation says it waits for two consecutive screenshots to match before comparing the result with the expected snapshot. That helps handle rendering that is still settling; it is not a queue retry policy for failed navigation. [Playwright visual comparisons]

2. Define a bounded, provider-specific policy

There is no universal list of retryable browser or screenshot API errors. Classify errors using the library or service you use. In general, retry a failure only when another attempt could plausibly succeed, such as a transient navigation timeout or temporary connection reset. Do not repeatedly retry malformed input, a persistent access denial, or an application page that consistently returns an error.

  • Set a small maximum number of attempts per URL.
  • Use exponential backoff with jitter so concurrent workers do not all retry together.
  • Record the attempt number, timestamp, error type, and message.
  • Give the URL a total time budget as well as a per-navigation timeout.
  • When the cap is reached, mark the item failed and continue with other URLs.

For example, AWS Batch job retry settings are scheduler-level: its guide documents between 1 and 10 attempts, and says timeout-terminated jobs are not retried under the described timeout guidance. Those semantics apply to AWS Batch; do not assume another scheduler behaves the same way. [AWS Batch retry strategy]

3. Runnable Python example with Playwright

Install the dependency and browser once:

python -m pip install playwright
python -m playwright install chromium

Save this as batch_capture.py. It writes each successful PNG atomically, persists a JSON manifest after every URL, resumes by skipping completed files, and retries only selected transient Playwright failures. The retry classifier is an example; review it against your browser version and network environment.

import json
import random
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import Error as PlaywrightError
from playwright.sync_api import sync_playwright

URLS = [
    "https://example.com/",
    "https://www.python.org/",
]
OUT = Path("captures")
MANIFEST = OUT / "manifest.json"
MAX_ATTEMPTS = 3
NAVIGATION_TIMEOUT_MS = 30_000

# Adapt this classifier to the errors surfaced by your provider/runtime.
TRANSIENT_ERROR_TEXT = (
    "net::ERR_CONNECTION_RESET",
    "net::ERR_CONNECTION_CLOSED",
    "net::ERR_TIMED_OUT",
    "net::ERR_NAME_NOT_RESOLVED",  # Often a DNS issue; retry is bounded.
)

def now():
    return datetime.now(timezone.utc).isoformat()

def load_manifest():
    if not MANIFEST.exists():
        return {}
    try:
        return json.loads(MANIFEST.read_text(encoding="utf-8"))
    except (json.JSONDecodeError, OSError):
        # Preserve a damaged manifest for inspection rather than silently
        # treating every prior success as lost.
        backup = MANIFEST.with_suffix(".json.corrupt")
        MANIFEST.replace(backup)
        return {}

def save_manifest(items):
    temp = MANIFEST.with_suffix(".json.tmp")
    temp.write_text(json.dumps(items, indent=2, sort_keys=True), encoding="utf-8")
    temp.replace(MANIFEST)

def valid_http_url(value):
    parsed = urlparse(value)
    return parsed.scheme in ("http", "https") and bool(parsed.netloc)

def is_retryable(exc):
    if isinstance(exc, PlaywrightTimeoutError):
        return True
    if isinstance(exc, PlaywrightError):
        message = str(exc)
        return any(marker in message for marker in TRANSIENT_ERROR_TEXT)
    return False

def capture_one(browser, url, item):
    for attempt in range(1, MAX_ATTEMPTS + 1):
        item.update(status="running", attempt_count=attempt, last_error=None,
                    updated_at=now())
        save_manifest(manifest)
        page = browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
        try:
            response = page.goto(url, wait_until="load", timeout=NAVIGATION_TIMEOUT_MS)
            # HTTP error responses may still render a useful page and produce
            # a screenshot. Record the status so the caller can decide whether
            # it counts as a successful capture.
            http_status = response.status if response else None
            page.screenshot(path=str(OUT / (item["file"] + ".tmp.png")), full_page=True)
            temp = OUT / (item["file"] + ".tmp.png")
            final = OUT / item["file"]
            temp.replace(final)
            item.update(status="succeeded", http_status=http_status,
                        completed_at=now(), updated_at=now(), last_error=None)
            save_manifest(manifest)
            return
        except (PlaywrightTimeoutError, PlaywrightError) as exc:
            item.update(last_error={"type": type(exc).__name__, "message": str(exc)},
                        updated_at=now())
            save_manifest(manifest)
            if attempt >= MAX_ATTEMPTS or not is_retryable(exc):
                item.update(status="failed", updated_at=now())
                save_manifest(manifest)
                return
            # Full jitter in [0, min(20s, 1s * 2^(attempt-1))].
            delay = random.uniform(0, min(20.0, 2 ** (attempt - 1)))
            time.sleep(delay)
        finally:
            page.close()

OUT.mkdir(parents=True, exist_ok=True)
manifest = load_manifest()

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    try:
        for index, url in enumerate(URLS):
            if not valid_http_url(url):
                manifest[url] = {"status": "failed", "attempt_count": 0,
                                 "last_error": {"type": "InvalidURL", "message": "Expected an http or https URL"},
                                 "updated_at": now()}
                save_manifest(manifest)
                continue

            file_name = f"page-{index:04d}.png"
            existing = manifest.get(url, {})
            # A file and a succeeded manifest entry together are the checkpoint.
            if existing.get("status") == "succeeded" and (OUT / file_name).is_file():
                continue

            item = existing
            item.update(url=url, file=file_name, status="pending",
                        attempt_count=item.get("attempt_count", 0), updated_at=now())
            manifest[url] = item
            capture_one(browser, url, item)
    finally:
        browser.close()

save_manifest(manifest)
print(f"Manifest: {MANIFEST}")
print(json.dumps({"succeeded": sum(x.get("status") == "succeeded" for x in manifest.values()),
                  "failed": sum(x.get("status") == "failed" for x in manifest.values())}, indent=2))

Run it with python batch_capture.py. On rerun, URLs with both a successful manifest entry and an output file are skipped. Replace URLS with your input list and consider storing stable IDs for URLs if the input order can change; this example’s filenames are based on list position.

Important implementation details

  • HTTP status is not automatically a capture exception. A 404 page may be worth capturing. Decide whether your product treats an HTTP error page as a successful screenshot, and store its status separately.
  • Navigation completion is a choice. wait_until="load" can be slow or time out on pages with long-running resources. Using domcontentloaded may be appropriate for some targets, but page content may still be incomplete. A fixed delay or selector wait should match the page being captured.
  • Check checkpoint consistency. The code uses an atomic rename for each image and manifest file. For multiple processes or machines, use a transactional database or object store with conditional writes and a lock/lease per URL.
  • Do not use URL text as a filesystem path. Use a hash or stable internal ID for filenames in production to avoid unsafe characters and collisions.
  • Retries can duplicate side effects. Keep capture work read-only where possible. If navigation triggers state-changing actions, make those operations idempotent or isolate them.

4. cURL, Python, and Node.js with a screenshot API

If your screenshot service provides a batch endpoint, use its documented per-item results and retry only failed items. Do not infer endpoint limits, error codes, retryability, or billing behavior from another service. If it offers only one-URL requests, the same manifest approach applies: issue one request per URL, save successful responses immediately, and retry only classified transient failures.

cURL: capture one URL per request

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/ \
  -o example.png

Python: loop over pending URLs

import requests

url = "https://example.com/"
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": url},
    timeout=90,
)
r.raise_for_status()
with open("example.png", "wb") as f:
    f.write(r.content)

Node.js: capture one URL

const q = new URLSearchParams({
  access_key: process.env.SCREENSHOTNEO_API_KEY,
  url: 'https://example.com/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('example.png', bytes));

These snippets show a single request. To build a resilient API batch, wrap each request in a per-URL state record, save the response to a temporary file and rename it on success, then persist the result before moving on. Set a timeout supported by your HTTP client, cap retries, and inspect the service’s response headers and documentation to determine whether a response represents a clean screenshot, a failed load, or another outcome.

5. Checkpointing and scheduler retries

A manifest can be JSON for a small single-process batch. For larger or concurrent jobs, store records in a database or durable queue with fields such as:

{
  "url_id": "stable-input-id",
  "url": "https://example.com/",
  "status": "pending | running | succeeded | failed",
  "attempt_count": 0,
  "started_at": null,
  "completed_at": null,
  "last_error": null,
  "output_key": null
}

Claim work atomically so two workers do not capture the same URL at once. On process restart, return stale running records to pending only after checking whether an output was already committed. Make the output write and status transition recoverable: a crash between those operations should not lose a good image or mark a missing one complete.

If you also configure scheduler retries, ensure the job loads the manifest and skips completed URL records. AWS Batch documents that jobs are attempted once by default and can be configured for up to 10 attempts; its timeout behavior has a separate rule for timeout-terminated jobs. Check the documentation for your scheduler and avoid treating job-level retries as a substitute for item-level state. [AWS Batch retry strategy]

6. Preserve useful diagnostics

Emit a final manifest or report with one row per original URL: final status, attempts, output location, HTTP status if available, and the last error. Keep enough details to distinguish DNS resolution, navigation timeout, browser crash, access denial, screenshot encoding, and storage failure. Avoid logging credentials, cookies, authorization headers, or sensitive query parameters.

For Playwright Test, retry metadata is available through TestInfo, and test attachments can preserve screenshot files or content for later inspection. Attach diagnostics when they help explain a test failure. [Playwright TestInfo]

Keep the rendering environment stable when screenshots are compared visually. Playwright notes that results can vary by host operating system, browser version and settings, hardware, power source, and headless mode. Pin browser and operating system versions where practical, and generate baselines in the same environment used for comparisons. [Playwright visual comparisons]

7. Troubleshooting

Symptom Likely cause Fix
Successful URLs are captured again after restart Progress is only in memory, or success is recorded before output is durable Persist per-URL state and output atomically; skip only when both checkpoint and output are present
The same URL retries until the batch runs out of time No per-URL attempt cap or the scheduler keeps restarting the whole batch Set a local attempt cap and total time budget; checkpoint completions before scheduler retry
Every retry fails with the same navigation error The cause is persistent, such as invalid input or access denial Stop retrying that class of failure, retain evidence, and inspect the URL or access requirements
Capture times out on an otherwise visible page The chosen load condition waits for a resource that never settles Choose a suitable navigation condition or wait for a page-specific selector; keep a hard timeout
HTTP 404 or 500 is treated as success Navigation completed and produced a response; transport success differs from application success Record HTTP status independently and apply your own acceptance rule
Two workers overwrite or duplicate a capture No atomic claim, lease, or unique output key Use a database/queue claim and stable URL ID; make writes conditional or idempotent
Visual checks fail inconsistently on identical input Browser or host rendering conditions differ Pin the environment and inspect fonts, browser version, headless mode, and hardware conditions
A retry overwrites a good image with a partial file Output was written directly to the final path Write to a temporary path and rename only after capture completes

8. Performance, reliability, and cost

  • Concurrency: run a bounded number of captures in parallel. More workers can increase throughput, but also increase memory use, target-site load, rate limits, and synchronized failures. Add per-host limits when needed.
  • Backoff: exponential backoff with jitter reduces retry bursts. The exact base delay and cap should fit your total job deadline and provider limits.
  • Timeouts: budget for navigation, rendering, screenshot encoding, and storage. A scheduler timeout shorter than the per-URL budget can terminate work before item-level recovery runs.
  • Idempotency: use stable item IDs and deterministic output keys. Check for an already committed image before repeating expensive work.
  • Cost accounting: retries may consume browser time or API requests. Confirm your service’s billing and cache rules. Do not assume a failed capture is free unless the provider explicitly says so.
  • Observability: track success rate, retry count, error class, elapsed capture time, and final unresolved URLs. Avoid using raw URLs as metric labels if they create excessive cardinality or expose sensitive query data.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single request captures one URL; put the same request inside your per-URL retry and checkpoint loop, and consult the ScreenshotNeo API documentation for request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

With ScreenshotNeo, cookie and consent banners, newsletter popups, and chat widgets are removed before the screenshot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify page verdict and billing. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. All features are on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

10. FAQ

Should I retry a screenshot that returned an HTTP error status?

Not automatically. A browser may successfully capture an error page. Store the HTTP status separately and decide whether that page is acceptable for your use case.

How many retries should I allow?

Choose a small cap that fits the total job deadline and your provider’s guidance. There is no universal retry count for browser or screenshot failures.

Does Playwright’s screenshot assertion retry failed URLs?

No. Its visual assertion waits for consecutive matching captures before comparing a snapshot. Handle navigation and batch-item retries in your own workflow or test setup.

Can I safely retry the whole scheduled job?

Yes, if each completed URL is durably checkpointed and the restarted job skips those results. Otherwise, a job retry can repeat successful captures.

Sources