ScreenshotNeo

BlogEngineering

How Timeouts Work in Web Scraping APIs

Understand API deadlines, browser render waits, retries, status codes, and practical timeout settings for reliable web scraping.

By the ScreenshotNeo team1 October 20268 min read

Short answer: a scraping timeout is rarely one universal clock. Most systems have a client-side deadline, an API request deadline, and browser-rendering waits such as a fixed delay, a CSS selector, or a page-load condition. Configure each for the failure mode you are addressing, then inspect the response status and body before deciding whether to retry.

A timeout value always belongs to a specific provider and endpoint. Check its unit, default, minimum, maximum, billing behavior, and retry rules in the current API reference. The figures in this guide for ScrapingBee are ScrapingBee settings, not industry standards.

1. The timeout layers

Client-side deadline

Your HTTP client can stop waiting before the scraping provider finishes. This protects your application thread, worker, or serverless function. It does not necessarily cancel work already running at the provider.

Provider request timeout

The API applies its own upper bound to the request. Depending on the service, this can include queueing, browser startup, DNS and connection work, page loading, JavaScript execution, and response transfer. Other services apply the limit only to part of that pipeline.

Browser render and readiness waits

A rendered page may be returned before a JavaScript application has populated the element you need. Rendering controls solve that problem:

  • Fixed wait: pause for a specified duration.
  • Selector wait: continue when a CSS or XPath selector appears.
  • Browser event wait: continue after a documented browser condition such as load or network idle.

Increasing the overall API deadline does not guarantee that the required content will appear. If the page is dynamic, use the readiness condition that represents the content you need.

Transfer and downstream deadlines

After the provider finishes, your proxy, load balancer, runtime, or object-storage upload can impose another deadline. Set those limits consistently so an intermediary does not close a connection while the provider is still within its documented limit.

2. A documented provider example: ScrapingBee

ScrapingBee’s HTML API documents timeout in milliseconds. Its documented default is 140,000 ms, and the accepted range is 1,000–140,000 ms. ScrapingBee also states that changing the value could negatively affect success rate and documents a 0.5-second margin of error. These are vendor-specific details; verify the current reference before using them.

ScrapingBee separates the request deadline from render readiness:

  • wait accepts 0–35,000 ms for a fixed JavaScript render wait.
  • wait_for waits for a CSS or XPath selector.
  • wait_browser waits for a documented browser condition.

A practical configuration is to keep the provider deadline within its allowed range, then choose the smallest readiness wait that reliably exposes the data. A long fixed sleep often adds latency without proving that the page is ready.

3. How long should you wait?

Situation First setting to adjust Why
Static HTML is slow to download Provider request timeout The page may need more time for connection or transfer.
Content is inserted by JavaScript Selector or browser-condition wait Readiness is the problem, not only the total deadline.
Images or cards appear after scrolling Full-page or lazy-load behavior The browser must trigger loading before capture or extraction.
Only some targets are slow Bounded per-request timeout and retry policy Do not make every request pay for the slowest site.
Your worker exits first Client, proxy, and runtime deadlines The provider cannot return a result to a closed connection.

Use measurements from your own targets. Record total duration, HTTP status, provider error code, response body, and whether the target was a static or JavaScript-rendered page. Set a deadline that covers normal variance while retaining a bounded upper limit.

4. Runnable timeout handling

Python with a generic scraping endpoint

Replace SCRAPING_API_URL and parameter names with those documented by your provider. The timeout passed to requests is your client-side deadline; it is separate from the provider’s timeout parameter.

import time
import requests

API_URL = "https://SCRAPING_API_URL"
API_KEY = "YOUR_API_KEY"
TARGET = "https://example.com"

params = {
    "api_key": API_KEY,
    "url": TARGET,
    # Use the provider's documented unit and bounds.
    "timeout": 30000,
    # Enable only if the target needs JavaScript.
    "render_js": "true",
    # Prefer a readiness condition over an arbitrary long sleep.
    "wait_for": ".product-card",
}

try:
    response = requests.get(API_URL, params=params, timeout=(10, 90))
    print("provider status:", response.status_code)
    print("body:", response.text[:1000])
    response.raise_for_status()
except requests.Timeout:
    print("The client deadline expired before a response arrived")
except requests.RequestException as exc:
    print(f"Request failed: {exc}")

Node.js with bounded retries

const API_URL = 'https://SCRAPING_API_URL';
const baseParams = {
  api_key: 'YOUR_API_KEY',
  url: 'https://example.com',
  timeout: '30000',
  render_js: 'true',
  wait_for: '.product-card'
};

function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

async function scrape(attempts = 3) {
  for (let attempt = 1; attempt <= attempts; attempt++) {
    const controller = new AbortController();
    const timer = setTimeout(() => controller.abort(), 90000);
    try {
      const query = new URLSearchParams(baseParams);
      const response = await fetch(`${API_URL}?${query}`, {
        signal: controller.signal
      });
      const body = await response.text();
      if (response.ok) return { status: response.status, body };

      const retryable = response.status === 408 || response.status === 429 || response.status >= 500;
      if (!retryable || attempt === attempts) {
        throw new Error(`Non-retryable or exhausted response ${response.status}: ${body}`);
      }
    } catch (error) {
      if (attempt === attempts) throw error;
    } finally {
      clearTimeout(timer);
    }
    await sleep(1000 * 2 ** (attempt - 1));
  }
}

scrape().then(result => console.log(result.status, result.body.slice(0, 500)))
  .catch(error => console.error(error.message));

cURL request pattern

curl --fail-with-body --max-time 90 -G "https://SCRAPING_API_URL" \
  --data-urlencode "api_key=YOUR_API_KEY" \
  --data-urlencode "url=https://example.com" \
  --data-urlencode "timeout=30000" \
  --data-urlencode "render_js=true" \
  --data-urlencode "wait_for=.product-card"

--max-time is the cURL client deadline. It does not change the provider’s server-side timeout. Keep the provider’s documented parameter separate and use its required unit.

5. What happens when a scraping API times out?

  1. The client may receive a local timeout exception if its own deadline expires.
  2. The provider may return an HTTP error with a provider-specific code and message.
  3. The provider may return a response describing a target failure, even when the target itself used another status.
  4. A queued or running job may continue after a synchronous client disconnects, depending on the service.

Always inspect both the status and response body. ScrapingBee documents that its default mapping converts many target errors to provider-side HTTP 500 responses. Its transparent_status_code=true option changes that mapping, but also disables its retry behavior and has billing implications. Use it only after checking the current documentation and understanding the tradeoff.

6. Retries, backoff, and idempotency

Retry transient transport or provider failures; do not use retries to hide a deterministic selector mistake, authentication error, robots denial, or consistently unreachable target.

  • Use a small maximum attempt count.
  • Apply exponential backoff with jitter when many workers may retry together.
  • Retry only statuses and exceptions documented as transient by the provider.
  • Keep the operation idempotent, or attach a request identifier so duplicate work can be detected.
  • Log the final response body after retries are exhausted.

ScrapingBee’s CLI documentation describes three attempts by default for transient 5xx and connection errors, with documented delays of 2, 4, and 8 seconds. That behavior is specific to its CLI and should not be assumed for every API client. Oxylabs’ guide lists HTTP-like code 524 as timeout or service unavailable; error-code meanings vary by provider.

7. Troubleshooting checklist

Symptom Likely cause Fix
Client reports a timeout but the provider shows a completed job Client or proxy deadline is shorter than provider processing time Increase the client deadline, use asynchronous jobs, or poll a job endpoint.
HTML arrives but expected content is missing JavaScript has not finished rendering Enable rendering and wait for the required selector or browser condition.
Every request returns 500 Provider status mapping, authentication, or a target-side failure Read the response body and provider logs; do not assume the target returned 500.
Retries never help Deterministic error or invalid parameter Validate URL, credentials, selector syntax, parameter units, and documented bounds.
Only large pages fail Transfer, memory, or browser-rendering cost exceeds the deadline Capture or extract only the needed content, block unnecessary resources, or use an asynchronous workflow.
Timeout parameter is rejected Wrong unit or outside provider limits Check the endpoint reference for spelling, unit, minimum, maximum, and default.
Results vary between attempts Dynamic content, rate limits, consent overlays, or third-party scripts Wait for a stable selector, control headers or cookies where supported, and record the page verdict.

8. Performance, reliability, and cost

Performance

  • Use static fetching when JavaScript is unnecessary.
  • Prefer selector waits over long fixed delays.
  • Block ads, trackers, and unused resource types when the provider supports it.
  • Cache stable pages with a documented TTL.
  • Use bulk or asynchronous endpoints for large batches when available.

Reliability

  • Separate transient provider failures from target-specific failures.
  • Track latency percentiles and timeout rates by target domain.
  • Preserve response headers and bodies for diagnosis.
  • Set an overall workflow deadline in addition to each request deadline.
  • Use a dead-letter queue for targets that repeatedly exceed the limit.

Cost

Longer waits can consume more browser and proxy capacity, and retries can multiply usage. Check whether a provider bills failed, timed-out, cached, or retried requests. A timeout policy should balance the value of a complete result against the cost of waiting and retrying.

9. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It handles consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for the full option set, including selector and full-page capture, device and retina settings, custom CSS and JavaScript, waits, request blocking, cookies and headers, timezone and geolocation, caching, signed links, asynchronous jobs, webhooks, bulk capture, PDFs, and usage reporting. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Free accounts include 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

10. FAQ

Is a timeout the same as a failed scrape?

No. It means a configured deadline was reached. The target may still be available, or the content may simply require a different readiness condition.

Should I always use the maximum timeout?

No. Start with a bounded value appropriate for the target and workload. Maximum values can increase latency, tie up workers, and raise retry costs.

Which wait should I choose for a JavaScript page?

Use a selector when a specific element proves readiness. Use a browser event when the provider documents a reliable event for the page. Use a fixed wait only when the page has no dependable readiness signal.

Can I solve timeouts only by adding retries?

No. Retries help with transient failures. They do not fix an invalid selector, a blocked target, an underspecified render wait, or a client deadline that is too short.

Why did the API return status 500 when the website works in my browser?

The provider may map target errors to its own status, or the browser and scraper may receive different content. Read the provider response body and compare rendering, headers, cookies, and readiness settings.