ScreenshotNeo

BlogEngineering

Handling Network Failures in Screenshot APIs

Classify screenshot failures correctly, set bounded timeouts, retry safely, and capture diagnostics with practical Playwright, Puppeteer, cURL, Python, and Node.js patterns.

By the ScreenshotNeo team1 October 202611 min read

Screenshot jobs fail in several different ways, and the right fix depends on which stage failed. A DNS error, a browser timeout, an HTTP 503 response, and a page that loaded without its images should not share one generic retry rule.

The reliable pattern is:

  1. Set separate budgets for the whole job, navigation, actions, and individual requests.
  2. Classify transport failures separately from HTTP responses.
  3. Retry only idempotent, transient work, with jittered exponential backoff and a hard attempt limit.
  4. Capture evidence before retrying: URL, phase, elapsed time, status, error text, failed resources, and a correlation ID.
  5. Return a stable error contract to callers.

Playwright explicitly treats HTTP errors such as 404 and 503 as completed HTTP responses; they are not requestfailed events. A transport failure means the client could not obtain an HTTP response. Playwright Page API documentation

1. Classify the failure before choosing a fix

Class Examples Retry? Typical action
DNS or connection NXDOMAIN, connection refused, reset Usually, if transient Retry with backoff; check resolver and origin availability.
TLS Certificate validation, handshake failure Usually no Fix the certificate, hostname, or trust configuration.
Navigation timeout Document never reaches the required load state Sometimes Increase a bounded navigation budget, wait for a more suitable condition, or reduce blocking work.
Resource failure Image, font, script, or API request fails Per resource Record the failed URL; decide whether the screenshot is still usable.
HTTP response 404, 429, 500, 503 Depends on status Apply status policy. Do not confuse it with a transport exception.
Browser or worker crash Target closed, renderer crashed Yes for a fresh worker Discard the worker, preserve diagnostics, and retry within the job deadline.
Incomplete render HTML loaded but lazy images or fonts are missing Not automatically Wait for a selector, network idle, or an application-specific readiness signal.

HTTP status is evidence about the target server. A request failure is evidence that the browser or network could not complete an exchange. Log both independently.

2. Set bounded timeouts at every layer

Use an overall deadline shorter than your queue visibility timeout. Inside it, reserve time for connection setup, navigation, resources, rendering, screenshot encoding, and response delivery. A single unlimited wait lets one stalled asset consume the entire job.

Budget Purpose
Overall job deadline Hard upper bound returned to the API caller.
Connect budget DNS, TCP, and TLS establishment.
Navigation budget Initial document request and selected load state.
Resource budget Optional cap for individual requests or readiness resources.
Action budget Clicks, selector waits, and scripted interactions.
Render budget Final layout, lazy content, and image encoding.

Puppeteer wait methods document a 30-second default timeout, and Page.setDefaultTimeout() can change it. Set an explicit value rather than relying on a library default. Playwright operations accept timeout options and support cancellation with AbortSignal. Puppeteer timeout documentation Playwright navigation documentation

Playwright timeout and cancellation example

import { chromium } from 'playwright';

const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();

const jobDeadline = AbortSignal.timeout(45_000);
page.setDefaultTimeout(8_000);

try {
  const response = await page.goto('https://example.com', {
    waitUntil: 'domcontentloaded',
    timeout: 20_000,
    signal: jobDeadline
  });

  console.log({
    status: response?.status() ?? null,
    url: page.url()
  });

  await page.screenshot({ path: 'shot.png', fullPage: true, timeout: 10_000 });
} finally {
  await browser.close();
}

If your installed Playwright version does not expose a signal option on a particular operation, enforce the same deadline with a timer that closes the page or context, and make the cancellation path explicit in your job code.

Puppeteer timeout example

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();
page.setDefaultTimeout(8_000);

try {
  const response = await page.goto('https://example.com', {
    waitUntil: 'domcontentloaded',
    timeout: 20_000
  });
  console.log({ status: response?.status() ?? null });
  await page.screenshot({ path: 'shot.png', fullPage: true });
} finally {
  await browser.close();
}

3. Retry only safe, transient work

A screenshot GET or a browser navigation is normally idempotent: repeating it does not mutate the target site. That makes retries reasonable for transient DNS failures, connection resets, connection refusals, provider rate limits, and selected 5xx responses.

Do not retry invalid URLs, authentication failures, deterministic 4xx responses, certificate errors, or the same identical failure indefinitely. A retry cannot repair a malformed URL or an expired credential.

Playwright API request context documents maxRetries; its current built-in behavior retries ECONNRESET, while HTTP response codes are not retried automatically. Build your own status policy. Playwright APIRequestContext documentation

Jittered exponential backoff

function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

function isRetryableError(error) {
  return ['ECONNRESET', 'ECONNREFUSED', 'ENOTFOUND', 'ETIMEDOUT'].includes(error.code);
}

function isRetryableStatus(status) {
  return status === 408 || status === 425 || status === 429 || status >= 500;
}

async function withRetries(operation, { attempts = 3, baseDelayMs = 400 } = {}) {
  let lastError;
  for (let attempt = 1; attempt <= attempts; attempt++) {
    try {
      const result = await operation(attempt);
      if (result.status && isRetryableStatus(result.status) && attempt < attempts) {
        const jitter = Math.floor(Math.random() * 200);
        await sleep(baseDelayMs * (2 ** (attempt - 1)) + jitter);
        continue;
      }
      return result;
    } catch (error) {
      lastError = error;
      if (!isRetryableError(error) || attempt === attempts) throw error;
      const jitter = Math.floor(Math.random() * 200);
      await sleep(baseDelayMs * (2 ** (attempt - 1)) + jitter);
    }
  }
  throw lastError;
}

Keep the attempt limit small and enforce the overall deadline around the retry loop. Respect a provider’s Retry-After value when one is available, while still enforcing your own maximum wait.

4. Capture evidence before retrying

A retry without diagnostics hides the failure pattern. For every attempt, record:

  • Target URL and a normalized hostname.
  • Phase: DNS, connect, TLS, navigation, resource, rendering, or delivery.
  • Attempt number, start time, elapsed milliseconds, and remaining job time.
  • HTTP status and response headers when a response exists.
  • Browser, worker, and provider identifiers.
  • Error name, code, and message.
  • Failed request URLs and resource types.
  • Console errors, final HTML or a safe excerpt, and a diagnostic screenshot or trace when policy permits.

Playwright request and failure logging

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();
const failed = [];

page.on('requestfailed', request => {
  failed.push({
    url: request.url(),
    method: request.method(),
    resourceType: request.resourceType(),
    errorText: request.failure()?.errorText ?? 'unknown'
  });
});

page.on('response', response => {
  if (response.status() >= 400) {
    console.log({ url: response.url(), status: response.status() });
  }
});

page.on('console', message => {
  if (message.type() === 'error') console.log({ console: message.text() });
});

try {
  const response = await page.goto('https://example.com', {
    waitUntil: 'domcontentloaded',
    timeout: 20_000
  });
  await page.screenshot({ path: 'diagnostic.png', fullPage: true });
  console.log({ status: response?.status() ?? null, failed });
} finally {
  await browser.close();
}

Routing lets you reproduce resource-specific failures and verify whether blocking a known tracker or advertisement changes readiness. Keep request interception narrow: blocking a critical API, stylesheet, or font can create an incomplete screenshot that looks like a network outage.

5. Decide whether an HTTP error should produce an image

Many screenshot systems should return a diagnostic image for a 404 or 503 page, because the browser did receive and render that response. Others should reject non-2xx pages for monitoring workflows. Make the policy explicit:

  • Visual archive: accept any rendered document and include its status in metadata.
  • Availability check: mark 4xx and 5xx as failed even if a screenshot was produced.
  • Mixed policy: accept 404 for content inventory, retry 429 and selected 5xx responses, and stop on authentication or validation errors.

Never infer transport success from a 200 status on a subresource. The main document can be 200 while critical images, scripts, or API calls failed.

6. Handle incomplete pages and lazy content

domcontentloaded ends sooner than a full visual readiness condition. Choose a readiness signal that matches the site:

  1. Navigate with a bounded timeout.
  2. Wait for a stable application selector such as [data-screenshot-ready].
  3. For lazy images, scroll or use the framework’s readiness hook.
  4. Use a short render delay only when the page has no better signal.
  5. Capture and record failed resources.
await page.goto('https://example.com/dashboard', {
  waitUntil: 'domcontentloaded',
  timeout: 20_000
});
await page.waitForSelector('[data-screenshot-ready]', { timeout: 8_000 });
await page.screenshot({ path: 'dashboard.png', fullPage: true, timeout: 10_000 });

Network idle is useful for pages that finish loading through fetch calls, but it can be misleading on applications with polling, analytics, or long-lived connections. Prefer a domain-specific selector when possible.

7. Stable API error contracts

Callers should not parse browser exception strings. Return a stable object such as:

{
  "category": "navigation_timeout",
  "retryable": true,
  "attempts": 2,
  "elapsed_ms": 39120,
  "status": null,
  "phase": "navigation",
  "correlation_id": "job-7f2c"
}

Keep provider-specific details in logs, and expose enough information for a caller to decide whether to retry. A correlation ID should connect the API response, worker logs, trace, and stored diagnostic artifacts.

8. Practical fixes for common errors

Error Likely cause Fix
ENOTFOUND DNS lookup failed or hostname is wrong. Validate the URL, check resolver access, and retry only when the failure is intermittent.
ECONNREFUSED Nothing is accepting connections on the destination port. Check origin health and firewall rules; use bounded retries.
ECONNRESET Peer or intermediary closed the connection. Retry idempotent work with backoff and capture attempt data.
TLS or certificate error Hostname, certificate chain, or trust configuration is invalid. Fix certificate configuration. Do not hide the error by disabling verification in production.
Navigation timeout Slow origin, never-ending requests, or an unsuitable wait condition. Set separate budgets, use a readiness selector, and inspect pending resources.
HTTP 429 Target or provider rate limit. Honor rate-limit guidance, reduce concurrency, and retry with jitter.
HTTP 503 Service unavailable, but an HTTP response was received. Log the status as an HTTP result; retry only under your transient-status policy.
Blank screenshot Early capture, blocked critical resources, client-side error, or bot challenge. Capture console and failed-request evidence, wait for a real readiness signal, and inspect the final HTML.
Browser target closed Worker crash, OOM, or browser shutdown. Discard the worker, preserve diagnostics, and retry once within the job deadline.

9. Performance, reliability, and cost

Performance

  • Reuse a browser process when safe, but isolate contexts and close pages promptly.
  • Keep navigation, resource, and render budgets independent so one asset cannot consume all available time.
  • Block nonessential ads and trackers only after confirming they are not required for layout or readiness.
  • Use a selector or application readiness event instead of a long fixed delay.
  • Limit concurrency according to CPU, memory, and target rate limits; excess concurrency increases queueing and failure rates.

Reliability

  • Give each job a deadline shorter than queue visibility expiration.
  • Use idempotency keys or deduplication when the same URL can enter the queue repeatedly.
  • Separate retryable from terminal errors in metrics.
  • Track success by phase, status, hostname, and worker version.
  • Store a final diagnostic artifact for repeated failures, subject to privacy policy.

Cost

Retries consume browser time and provider request capacity. Bound attempts and backoff, and avoid retrying deterministic failures. If you use a hosted screenshot API, understand whether failed loads, cache hits, and HTTP error pages are billable before setting aggressive retry policies.

10. cURL, Python, and Node.js request patterns

These examples show a bounded client-side timeout. Your server should still enforce an overall job deadline and classify the response separately from transport errors.

cURL

curl --fail-with-body \
  --connect-timeout 10 \
  --max-time 45 \
  -o shot.webp \
  "https://your-screenshot-service.example/v1/shot?url=https%3A%2F%2Fexample.com"

Python

import requests

try:
    response = requests.get(
        "https://your-screenshot-service.example/v1/shot",
        params={"url": "https://example.com"},
        timeout=(10, 45),  # connect timeout, total read timeout
    )
    print("status:", response.status_code)
    response.raise_for_status()
    with open("shot.webp", "wb") as output:
        output.write(response.content)
except requests.exceptions.Timeout:
    print("transport timeout: retry only if the job is still within its deadline")
except requests.exceptions.RequestException as error:
    print("request error:", error)

Node.js

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 45_000);

try {
  const response = await fetch(
    'https://your-screenshot-service.example/v1/shot?url=https%3A%2F%2Fexample.com',
    { signal: controller.signal }
  );
  console.log('status:', response.status);
  if (!response.ok) throw new Error(`HTTP ${response.status}`);
  const bytes = Buffer.from(await response.arrayBuffer());
  await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
} catch (error) {
  console.error(error.name === 'AbortError' ? 'client timeout' : error);
} finally {
  clearTimeout(timer);
}

11. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF, while the service handles browser capture and reports the result with response headers.

Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; the response identifies the page verdict and billing state with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo documentation for request options and response details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture with lazy images loaded, element capture by CSS selector, dark mode, device presets and custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector and network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.

Plans include 1,000 screenshots each month free with no card; paid plans start at $5 for 3,000 screenshots. The other plans are $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan.

Start with 1,000 free screenshots a month—no card required.

12. Checklist for production

  • Overall deadline is shorter than queue visibility timeout.
  • Connect, navigation, action, resource, and render budgets are explicit.
  • HTTP responses and transport failures are separate categories.
  • Retry policy covers only idempotent transient work.
  • Backoff includes jitter and respects rate limits.
  • Every attempt records phase, status, elapsed time, and correlation ID.
  • Failed requests and console errors are retained for diagnosis.
  • Readiness uses a selector or application signal where possible.
  • Concurrency is limited and measured.
  • Clients receive a stable error contract with a retryable flag.

FAQ

Should every 503 be retried?

No. A 503 is an HTTP response, so the target did answer. Retry when your policy identifies it as transient, within a bounded deadline; preserve the response for monitoring and diagnosis.

What is the difference between a timeout and a failed request?

A timeout means an operation exceeded its configured budget. A failed request means the browser could not obtain an HTTP response for that request. A page can also return 503 successfully from the HTTP layer.

How many attempts should a screenshot job make?

Use a small hard limit such as two or three attempts, then stop when the overall deadline or repeated identical failure indicates the problem is deterministic.

Is network idle always the best readiness condition?

No. Polling and analytics can prevent idle, while a page can become visually ready before idle. A stable application selector is usually more precise.

Can a screenshot be useful when some resources failed?

Yes. Decide based on the page’s purpose. Record the failed resources and expose an incomplete-assets signal so consumers do not mistake a partial render for a perfect capture.