ScreenshotNeo

BlogHow-to

How to Retry a Scheduled Website Screenshot When the Capture Fails

Build bounded retries for scheduled website screenshots, distinguish capture failures from missed triggers, and keep the last good image safe.

By the ScreenshotNeo team4 October 202611 min read

Short answer: Put the retry around the entire scheduled capture job, at the script, workflow, or scheduler layer. Retry only plausible transient failures, cap the number of attempts, log each failure, and write each attempt to a temporary file. Mark the run successful only after the expected screenshot exists and is nonempty. Playwright’s screenshot assertion retry is for visual tests; it does not rerun a failed cron job or scheduled workflow.

This guide uses Playwright with Node.js for the browser capture, plus a bounded retry wrapper and a GitHub Actions schedule example. The same recovery principles apply if a different scheduler launches your capture script.

1. Separate the scheduled job from the screenshot

A scheduled screenshot has two distinct operations:

  1. The scheduler starts a job. A cron service or hosted workflow decides when the command should run.
  2. The capture job visits the page and writes an image. Browser launch, navigation, page readiness, screenshot generation, and file writing all happen here.

A retry at the wrong layer will not fix the failure. If no run started, inspect the scheduler and its run history. If the job ran but the browser failed, retry the capture operation. If the browser created an image but artifact upload failed, retry or repair the upload step according to that step’s behavior.

Playwright’s toHaveScreenshot() waits for two consecutive page screenshots to match, then compares the last screenshot with the expected baseline. Its timeout controls how long that assertion can retry while waiting for visual stability. It is not a policy for rerunning a scheduled job. [Playwright PageAssertions documentation]

2. Inspect the failed run before retrying

Open the scheduler’s run history and logs first. Confirm that the scheduled trigger started, then find the failing step and preserve its original error text. Record the scheduled occurrence, attempt number, and output path so repeated failures can be diagnosed.

Classify the failure before deciding whether another attempt can help:

Failure type Examples Typical response
Trigger or runner No workflow run, unavailable runner, queued job dropped Check scheduler history and runner availability; a capture retry cannot rerun a trigger that never started.
Browser startup Browser executable missing, launch process exits Check installed dependencies and runtime image. Retry only if evidence suggests a temporary startup issue.
Navigation or readiness Connection reset, navigation timeout, selector never appears Retry transient network or upstream failures; fix an incorrect selector or URL.
Screenshot generation Invalid selector, browser crash during capture Inspect the page and selector; retry only if the crash may be transient.
Write or upload Permission denied, disk full, artifact service error Fix permissions or storage capacity; retry a transient upload failure if the image was safely preserved.

These categories are a practical diagnostic aid, not a taxonomy imposed by Playwright. Do not retry invalid URLs, invalid credentials, missing dependencies, deterministic selector errors, or write-permission problems until their cause is corrected.

3. Add bounded retries around the complete capture

The following Node.js script launches Chromium, visits the page, waits for a chosen readiness condition, captures a full-page PNG, validates the temporary file, and then renames it to the final path. It retries the whole operation at most three times with increasing delays. Each attempt gets a separate temporary filename, so a failed attempt cannot overwrite the last successful screenshot.

import { chromium } from 'playwright';
import { mkdir, rename, stat } from 'node:fs/promises';
import path from 'node:path';

const url = process.env.TARGET_URL ?? 'https://example.com';
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const occurrence = process.env.OCCURRENCE ?? new Date().toISOString().replaceAll(':', '-');
const maxAttempts = 3;
const delaysMs = [1000, 3000];

function isLikelyTransient(error) {
  const message = String(error?.message ?? error);
  return /timeout|ECONNRESET|ECONNREFUSED|ENOTFOUND|socket|net::ERR_|Target closed|browser has disconnected/i.test(message);
}

async function capture(attempt) {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
    // Replace this with a stable, page-specific condition when available.
    await page.locator('body').waitFor({ state: 'visible', timeout: 10000 });

    const finalPath = path.join(outputDir, `${occurrence}.png`);
    const tempPath = path.join(outputDir, `${occurrence}.attempt-${attempt}.tmp.png`);
    await page.screenshot({ path: tempPath, fullPage: true, animations: 'disabled' });
    const result = await stat(tempPath);
    if (result.size === 0) throw new Error('Screenshot file is empty');
    await rename(tempPath, finalPath);
    return finalPath;
  } finally {
    await browser.close();
  }
}

await mkdir(outputDir, { recursive: true });
let lastError;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
  try {
    const file = await capture(attempt);
    console.log(JSON.stringify({ status: 'success', file, attempt, url }));
    process.exit(0);
  } catch (error) {
    lastError = error;
    console.error(JSON.stringify({ status: 'attempt_failed', attempt, maxAttempts, url, error: String(error?.stack ?? error) }));
    if (attempt === maxAttempts || !isLikelyTransient(error)) break;
    await new Promise(resolve => setTimeout(resolve, delaysMs[attempt - 1] ?? 5000));
  }
}
console.error(JSON.stringify({ status: 'failed', attempts: maxAttempts, url, error: String(lastError?.stack ?? lastError) }));
process.exit(1);

Install the dependency and browser before scheduling the script: npm install playwright and npx playwright install chromium. Run it with TARGET_URL=https://example.com node capture.mjs. Provide OUTPUT_DIR and a stable OCCURRENCE value from your scheduler if you need predictable artifact names.

The transient-error matcher is a starting point, not a universal classifier. Adapt it to your runtime and the error types you see in logs. The script exits nonzero after a permanent error or exhausted attempts so the scheduler can report failure and alert you.

Choose page readiness deliberately

domcontentloaded is a useful starting point for pages that continue loading analytics or long-lived requests. It does not guarantee that every image or client-rendered section is ready. For a page with a clear landmark, wait for that landmark:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.locator('[data-capture-ready="true"]').waitFor({ state: 'visible', timeout: 15000 });

Use networkidle only when the page eventually becomes quiet; analytics, polling, and streaming connections can keep it from settling. Prefer a page-specific selector when the content matters. A fixed delay can help with a known animation or late content, but it adds time to every run and may still be too short during a slow response.

4. Schedule the script with GitHub Actions

This workflow runs the capture script on a cron schedule and on manual dispatch. It uploads the screenshot directory even after a failed capture, which helps retain available output and diagnostics. Replace the cron expression and URL with your schedule and target. Set TARGET_URL as a repository variable or workflow environment value; store credentials, if needed, as secrets.

name: Scheduled website screenshot

on:
  schedule:
    - cron: '17 * * * *'
  workflow_dispatch:

concurrency:
  group: website-screenshot-${{ github.ref }}
  cancel-in-progress: false

jobs:
  capture:
    runs-on: ubuntu-latest
    timeout-minutes: 10
    env:
      TARGET_URL: ${{ vars.TARGET_URL }}
      OCCURRENCE: ${{ github.run_id }}-${{ github.run_attempt }}
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - name: Capture page
        run: node capture.mjs
      - name: Upload screenshot artifacts
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: website-screenshot-${{ github.run_id }}-${{ github.run_attempt }}
          path: screenshots/
          if-no-files-found: ignore

The script already retries capture attempts. GitHub Actions can also rerun failed jobs manually or through its workflow controls, but be careful not to stack a large job-level retry on top of script retries: three capture attempts inside several whole-job reruns can multiply browser launches and runtime.

Scheduled GitHub Actions events may be delayed during high load, and some queued jobs may be dropped. If every scheduled capture matters, monitor whether each expected run occurred instead of relying on the cron declaration alone. [GitHub Actions schedule event documentation]

5. Prevent collisions and choose a concurrency policy

If a run takes longer than the schedule interval, the next occurrence may begin while the previous one is still working. Decide whether to allow parallel captures, cancel older runs, or queue work. Use occurrence-specific output names so concurrent runs do not write to the same temporary file.

GitHub Actions workflow runs can execute concurrently by default. A concurrency group limits active runs, but by default a newly pending run can replace an existing pending run. GitHub documents a queue option for workflows where pending runs should wait. Choose based on whether a newer screenshot may supersede an older pending capture or every occurrence must be retained. [GitHub Actions concurrency documentation]

For a local cron job, use a lock mechanism if overlapping runs would corrupt output or exceed resource limits. For a queue-based scheduler, make the queue and retention behavior explicit. Keep the last known-good screenshot until a new output has been validated and promoted.

6. Distinguish a new capture from a visual assertion retry

Use a job-level retry when the browser operation itself failed and another attempt may succeed. Use Playwright’s toHaveScreenshot() when a test must wait for a page to become visually stable before comparing it to a baseline. The assertion timeout is elapsed retry time for that assertion; it does not specify how many scheduled jobs to rerun. [Playwright PageAssertions documentation]

Example assertion in a Playwright Test test file:

import { test, expect } from '@playwright/test';

test('homepage matches its screenshot', async ({ page }) => {
  await page.goto('https://example.com');
  await expect(page).toHaveScreenshot('homepage.png', {
    fullPage: true,
    timeout: 15000
  });
});

This assertion requires a configured baseline and Playwright Test project. It is useful for visual regression checks, not as a replacement for scheduler-level failure handling.

7. Troubleshoot common failures

Symptom Likely cause Fix
No scheduled run appears Trigger delay/drop, disabled workflow, or schedule configuration issue Inspect Actions history and workflow configuration; add monitoring for expected run times. A browser retry cannot recover a trigger that never ran.
Executable doesn't exist or browser launch fails Chromium was not installed in the environment, or required system libraries are missing Install Playwright’s browser and dependencies in the same runner image used by the job.
Navigation timeout Slow site, transient network issue, or unsuitable readiness condition Check logs and target availability; use a realistic timeout and a page-specific readiness selector. Retry only if the cause appears transient.
Selector timeout Selector is wrong, content is conditional, or the page changed Inspect the page structure and update the selector or wait condition. Repeatedly retrying a deterministic mismatch hides the real issue.
Image is blank or incomplete Capture occurred before client rendering or lazy content completed Wait for a relevant element, scroll/load content if needed, or use a longer page-specific readiness condition. Validate output beyond file size if image correctness is critical.
Temporary image exists but final image is missing Rename or write failed, or output directory permissions are wrong Check runner permissions and disk space. Preserve the temporary file and error log for diagnosis.
Artifacts are missing after a failed run Upload step did not run or found no files Use an unconditional upload step and inspect its log; do not report capture success merely because upload ran.
Two runs overwrite each other Shared output path or overlapping schedule Include occurrence and attempt identifiers in temporary names, then set an explicit concurrency or queue policy.
Images differ across runs without a site change Browser, OS, hardware, settings, or headless-mode differences can affect rendering Keep the browser and execution environment consistent when comparing screenshots. Playwright documents these rendering sources of variation. [Playwright visual comparisons documentation]

8. Performance, reliability, and cost

  • Bound the total runtime. A rough upper bound is the per-attempt browser and navigation timeouts multiplied by the maximum attempts, plus the retry delays. Set the scheduler’s overall job timeout above the expected bound, but low enough to stop stuck jobs.
  • Use backoff sparingly. Increasing delays reduce rapid repeated load on a site or runner. Keep them short enough that retries finish before the next schedule when possible.
  • Avoid multiplying retries. If both the script and scheduler retry, calculate the maximum total attempts. Pick one layer for normal transient recovery and keep any outer rerun policy small.
  • Keep diagnostics. Log the attempt number, target, elapsed time, and error. Save a trace or other browser diagnostics when a failure is hard to reproduce, while avoiding secrets in logs and artifacts.
  • Control browser resource use. Close the browser in a finally block, keep concurrency within runner capacity, and avoid capturing more page area or more URLs than needed.
  • Budget for actual executions. Self-managed capture cost includes runner time, browser installation, storage, and any alerting or artifact retention. Retries consume additional compute, so cap them and monitor repeated failures.

For visual comparisons, keep execution conditions consistent where possible. Operating system, browser version, settings, hardware, power source, and headless mode can change rendering, so not every pixel difference indicates a website change. [Playwright visual comparisons documentation]

9. Or skip the browser setup

If you want a managed screenshot call, ScreenshotNeo is a website screenshot API and MCP server. You still run the schedule and decide whether to retry; the API handles the capture request. Its request options and response details are documented at ScreenshotNeo docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie banners and consent prompts are accepted, and known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan.

Sign up free for 1,000 screenshots a month, with no card required.

10. FAQ

How many times should a scheduled screenshot retry?

There is no universal count in the reviewed Playwright documentation. Start with a small cap, such as the three attempts in the example, then tune it to your schedule interval, runtime, and failure logs.

Should I retry every screenshot error?

No. Retry plausible transient runtime failures. Correct configuration problems such as invalid URLs, missing browsers, incorrect selectors, and permissions before running again.

Does Playwright automatically rerun a failed GitHub Actions schedule?

Playwright handles browser automation and visual assertions. The scheduler or workflow must implement scheduled-job reruns, and the scheduled trigger itself may not have run.

Can a retry replace the last successful screenshot?

It can if the workflow writes directly to the same path. Use a unique temporary path, validate the output, and promote it only after success.