How to Identify and Fix Timeouts in a Bulk Website Screenshot Job
Find which stage timed out before changing settings. Diagnose browser, caller, provider, queue, and capture failures, then apply targeted fixes.
A timeout tells you that a particular operation or caller deadline ran out; it does not prove that the site was unreachable or that no screenshot was produced. In a bulk screenshot job, first identify whether the deadline belonged to your client, browser navigation, page-readiness wait, screenshot operation, hosted provider, or render queue. Then change the setting at that boundary, record each URL independently, and check for an existing result before retrying.
This guide uses Playwright examples and covers hosted API and queue cases. Playwright defaults can vary by API language and configuration, so check the documentation for your installed version before relying on a default value. [Playwright Page API]
1. Log each URL and find the timeout boundary
Do not log only “batch failed.” Give every URL an independently inspectable outcome. Record:
- Batch or job ID, URL, attempt number, and timestamps.
- Elapsed time and the stage being run: navigation, readiness check, screenshot, upload, or provider request.
- HTTP status and response headers, or the browser exception and operation that threw.
- Whether an output file or provider-side result exists.
- Retry outcome and any provider request ID or job ID.
Compare failures with successes. If only slower sites fail, investigate navigation and page readiness. If failures rise as concurrency rises, inspect worker capacity and queueing. A 429 or 503 response points to rate limiting or temporary service availability. If the caller reports a timeout but an artifact exists, verify that result before capturing again. These are hypotheses to check against logs, not guaranteed diagnoses.
| Boundary | What to inspect | Typical next step |
|---|---|---|
| Caller/client deadline | HTTP client timeout, job deadline, proxy deadline | Check provider outcome or artifact; align client wait with the operation’s expected duration. |
| Browser navigation | The navigation call, wait condition, and browser exception | Choose a completion condition that matches the capture and verify needed content. |
| Page readiness | Whether the target element or data appeared | Wait for an explicit condition rather than an arbitrary extension. |
| Screenshot/output | Whether navigation finished and the capture or file handling failed | Separate capture errors from navigation and output-write errors. |
| Provider or queue | HTTP status, provider job state, queue delay, active slots | Use provider-specific status, retry, and capacity guidance. |
2. Fix browser navigation and readiness waits
Playwright navigation offers load, domcontentloaded, networkidle, and commit wait conditions. The documented default is load; timeout defaults and configuration can differ across APIs and language bindings. Record the exact operation that threw, and use a condition appropriate to the screenshot rather than extending every timeout. [Playwright Page API]
For example, a page with analytics, chat, polling, or streaming connections may never become network-idle. Playwright explicitly discourages networkidle as a testing readiness strategy and recommends web assertions to assess readiness. Wait for the content you actually need instead. Most Playwright actions auto-wait; add a separate wait only for a concrete condition. [Playwright Page API]
import { chromium } from 'playwright';
const urls = [
'https://example.com',
'https://example.org',
];
const browser = await chromium.launch();
const results = [];
for (const [index, url] of urls.entries()) {
const page = await browser.newPage();
const startedAt = new Date().toISOString();
try {
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
// Replace this with a selector that proves the content you need is ready.
await page.locator('body').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: `shot-${index}.png`, fullPage: true, timeout: 20_000 });
results.push({ url, startedAt, finishedAt: new Date().toISOString(), status: response?.status(), ok: true });
} catch (error) {
results.push({ url, startedAt, finishedAt: new Date().toISOString(), ok: false, error: String(error) });
} finally {
await page.close();
}
}
await browser.close();
console.log(JSON.stringify(results, null, 2));
This sequential example favors clear per-URL outcomes. To run concurrently, use a bounded worker pool and preserve the per-URL try/catch and result record. Avoid launching every URL at once: simultaneous browser contexts consume resources and may overload either your own workers or target sites.
3. Separate navigation, screenshot, and output handling
A successful navigation does not prove that screenshot capture or file writing succeeded. Keep these stages separately visible in logs. Playwright supports viewport and full-page screenshots and image format options; choose the smallest capture that meets the use case. Full-page captures of long or complex pages can take longer and produce larger artifacts. [Playwright screenshots]
If navigation succeeds but capture fails, reduce capture complexity where acceptable, check disk or object-storage writes, and inspect the screenshot operation’s own error. If the image exists but the job is marked failed, examine the upload, post-processing, or acknowledgment stage before recapturing.
4. Handle hosted API timeouts and retries safely
A client can stop waiting after the provider has already completed a capture. On timeout, check the provider’s job status, request record, or expected artifact before retrying. Otherwise a second successful capture may duplicate work or usage. Provider semantics vary, so consult the service’s documentation.
For transient 429 rate-limit or 503 service-unavailable responses, honor Retry-After when present. If it is absent, use capped exponential backoff with jitter and a bounded attempt count. Do not automatically retry malformed requests, invalid credentials, or exhausted quota; fix the cause first. ScreenshotEngine gives at most three retries as an example, not a universal retry rule. [ScreenshotEngine documentation]
import random
import time
import requests
TRANSIENT = {429, 503}
MAX_ATTEMPTS = 3
for attempt in range(MAX_ATTEMPTS):
try:
response = requests.get('https://provider.example/capture', timeout=(5, 60))
except requests.Timeout:
# Check provider job state or artifact before deciding to resubmit.
raise
if response.status_code not in TRANSIENT:
response.raise_for_status()
break
if attempt == MAX_ATTEMPTS - 1:
response.raise_for_status()
retry_after = response.headers.get('Retry-After')
delay = float(retry_after) if retry_after and retry_after.isdigit() else min(30, 2 ** attempt) + random.random()
time.sleep(delay)
The endpoint above is deliberately a placeholder: use the endpoint and response handling documented by your provider. A timeout exception does not expose whether remote work completed unless the provider offers a status mechanism or idempotency behavior.
5. Reduce queue and batch pressure
A render service can be healthy for an individual URL while requests time out waiting for a free renderer. Splash 3.1 documents that its timeout clock begins when a request arrives, so internal queue time consumes the request budget; it also notes that overloaded Splash can produce 504 errors. Track queue depth and, if available, distinguish request arrival from render start. [Splash 3.1 documentation]
For a large multi-site job, submit independently tracked work per URL and manage it as a batch. Splash’s documentation uses 100 websites as an example to split into per-site requests; that is an example, not a universal batch limit. Keep concurrency bounded and adjust it using observed queue delay, memory, CPU, and failure rates. There is no universal safe concurrency number.
If your deployment uses Splash, its documentation recommends resource timeouts so slow remote resources do not wait indefinitely and describes multiple instances behind a load-balancer-managed queue for capacity. Treat those as Splash-specific deployment recommendations. Increasing a request timeout without addressing a saturated queue can simply make callers wait longer.
6. cURL, Python, and Node.js examples for a bulk API workflow
For a provider that accepts one URL per request, a shell loop gives each URL an independent result. Adapt the endpoint, authentication, status handling, and output type to the actual API:
while IFS= read -r url; do
[ -z "$url" ] && continue
name=$(printf '%s' "$url" | sed 's#https\?://##; s#[^A-Za-z0-9._-]#_#g')
curl --fail-with-body --show-error --silent \
--connect-timeout 10 --max-time 90 \
-G 'https://provider.example/capture' \
--data-urlencode "url=$url" \
-o "$name.png" || printf 'FAILED %s\n' "$url" >&2
done < urls.txt
Python can isolate failures and keep going through the input list:
from pathlib import Path
from urllib.parse import urlparse
import re
import requests
session = requests.Session()
for url in Path('urls.txt').read_text().splitlines():
if not url.strip():
continue
safe_name = re.sub(r'[^A-Za-z0-9._-]+', '_', urlparse(url).netloc + urlparse(url).path).strip('_') or 'page'
try:
response = session.get(
'https://provider.example/capture',
params={'url': url, 'api_key': 'YOUR_API_KEY'},
timeout=(10, 90),
)
response.raise_for_status()
Path(f'{safe_name}.png').write_bytes(response.content)
print({'url': url, 'ok': True, 'status': response.status_code})
except requests.RequestException as exc:
print({'url': url, 'ok': False, 'error': str(exc)})
Node.js with built-in fetch can apply a caller-side deadline using AbortSignal.timeout in supported runtimes. That deadline only controls how long the client waits; check provider status before repeating a request:
import { writeFile } from 'node:fs/promises';
const urls = ['https://example.com', 'https://example.org'];
for (const [index, url] of urls.entries()) {
const query = new URLSearchParams({ url, api_key: 'YOUR_API_KEY' });
try {
const response = await fetch(`https://provider.example/capture?${query}`, {
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
await writeFile(`shot-${index}.png`, Buffer.from(await response.arrayBuffer()));
console.log({ url, ok: true, status: response.status });
} catch (error) {
console.error({ url, ok: false, error: String(error) });
}
}
Replace placeholder provider URLs and parameter names with the actual service’s documented interface. Never put a secret API key in a public client-side page.
7. Troubleshooting common timeout symptoms
| Symptom | Likely cause to investigate | Fix |
|---|---|---|
| Only a few slow sites time out | Navigation condition or required page content takes longer | Record the throwing operation; select the appropriate navigation condition and wait for a specific required element. |
| Pages never reach network idle | Persistent connections, analytics, polling, chat, or streaming | Avoid network-idle as a generic readiness test; assert on the content needed for the capture. |
| Failures increase with batch size or concurrency | Queue delay or worker resource pressure | Bound concurrency, split work per URL, and inspect queue and resource metrics. |
| HTTP 429 | Rate limit | Honor Retry-After, reduce request rate, and check provider limits. |
| HTTP 503 or renderer 504 | Transient service unavailability or overloaded renderer | Use bounded retry for transient errors and investigate provider capacity or queueing. |
| Client timeout but screenshot exists | Caller stopped waiting after remote success | Check job state or artifact before retrying. |
| Navigation succeeds but no image is saved | Screenshot, filesystem, upload, or post-processing stage failed | Log capture and output stages separately; verify destination permissions and available storage. |
| Every URL fails immediately | Invalid credentials, request format, DNS/network setup, or exhausted quota | Inspect the response body and provider status; correct configuration instead of retrying unchanged input. |
8. Performance, reliability, and cost
- Use bounded concurrency. Increase workers only while throughput improves without growing queue delay or failure rates.
- Choose the needed capture. Viewport screenshots are generally less work than full-page captures; avoid unnecessary image dimensions and post-processing.
- Give each URL its own state. Persist completed results so a later failure does not force a full batch restart.
- Make retries selective. Retry transient conditions with backoff; avoid repeating permanent failures or ambiguous successful captures.
- Separate time budgets. Set connection, navigation/readiness, capture, provider, and overall job deadlines deliberately. Leave the caller enough time to receive and persist the result.
- Measure before scaling. Track per-stage duration, queue wait, memory, CPU, and response status. A longer deadline can mask overload without increasing capacity.
Self-hosting gives direct control over browser and worker settings but requires operating capacity and queueing. A hosted API shifts browser operations to a provider, while adding provider-specific limits, response semantics, and usage rules. Compare timeout controls, queue visibility, concurrency limits, retry behavior, and duplicate-capture billing before choosing. No universal cost or concurrency figure applies across services.
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL as PNG, JPEG, WebP, or PDF. Its 63 options include full-page capture with lazy images loaded, selector capture, device and viewport settings, waits, custom headers and cookies, caching, async jobs, and bulk capture of up to 100 URLs per call. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000, and every feature is on every plan. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.
10. FAQ
Should I just increase the timeout?
Only after identifying the stage that expires. A longer deadline can help a genuinely slow operation, but it can also hide an overloaded queue or an unsuitable readiness condition.
Can a timeout happen after the screenshot succeeded?
Yes. The caller may stop waiting after remote capture completes. Check for a result or provider job state before resubmitting.
Is there a safe concurrency number for every screenshot job?
No. Capacity depends on the renderer, page mix, capture size, and environment. Bound concurrency and adjust it using measured queue delay and resource use.
Does networkidle mean the screenshot is ready?
Not reliably. Persistent connections may prevent it, and Playwright recommends assertions tied to the page content for readiness checks.
Sources
- Playwright Page API and screenshots guide.
- ScreenshotEngine documentation for its retry and timeout guidance.
- Splash 3.1 documentation for queue timeout behavior and renderer operations.


