ScreenshotNeo

BlogHow-to

PageCrawl.io API rate limits: how to handle 429 errors

A 429 means PageCrawl is rate-limiting your requests. Learn how to honor Retry-After, add backoff, pace calls, and check the limit for your endpoint and account.

By the ScreenshotNeo team4 October 20269 min read

HTTP 429 means PageCrawl.io is rate-limiting requests to an endpoint. Stop sending requests at the same pace, honor the Retry-After response header when it is present, and use backoff for automated retries. Do not retry immediately in a tight loop.

The applicable request ceiling depends on the endpoint and account. PageCrawl’s developer guide says its REST API allows 60 requests per minute on Free and 300 per minute on paid plans, while a separate dashboard guide describes 60 requests per minute as typical for most accounts. Check the current [PageCrawl API reference](https://pagecrawl.io/developers) for the endpoint and account you are using rather than treating either figure as universal.

1. What a 429 response tells you

A 429 response indicates that the server is limiting the request rate. By itself, it does not establish that the API is down. The Push API documentation instructs clients to honor Retry-After before retrying: “429 | Rate limited; honor the Retry-After header before retrying”.

PageCrawl documents different request limits in different guides. Its API developer guide lists 60 requests per minute for Free and 300 per minute for paid plans; its dashboard guide says most accounts have a 60-requests-per-minute limit. These are vendor-documented figures, not independent measurements. The API reference is described by PageCrawl as generated from its OpenAPI specification and taking precedence over guide text.

2. Respond to 429 in order

  1. Confirm the status. Record the request URL or endpoint, timestamp, account and plan context, response body, and response headers. Make sure the response is actually 429.
  2. Honor Retry-After. If supplied, wait as directed before retrying. Do not substitute a guessed fixed delay for the server’s instruction.
  3. Stop the current request burst. Pause or slow the producer of requests so new work does not immediately create another burst while retries are pending.
  4. Retry with backoff. If no Retry-After header is present, use a backoff policy that spaces retries and avoids a synchronized tight loop. PageCrawl’s dashboard guide recommends exponential backoff but does not establish a universal delay to use.
  5. Reduce unnecessary calls. Check for duplicate work and pace requests below the verified endpoint/account limit.
  6. Verify the applicable limit. Consult the current API reference for the endpoint and your account. If the published information does not resolve your situation, ask PageCrawl for confirmation; the reviewed documentation does not establish a process for requesting a higher limit.

3. Runnable retry examples

The examples below demonstrate the retry pattern for an authenticated PageCrawl request. Replace the illustrative endpoint and payload with the endpoint and request format from the current API reference. Do not assume every endpoint uses the same method or payload. PageCrawl documents Bearer token authentication for its authenticated Push API endpoints.

cURL: inspect a response and its headers

curl -i \
  -H "Authorization: Bearer $PAGECRAWL_API_TOKEN" \
  "https://pagecrawl.io/api/your-endpoint"

Use the actual endpoint shown in the API reference. The -i option includes response headers, so you can inspect the status and Retry-After. cURL does not implement a general policy here; check the header and wait before issuing another request.

Python: wait for Retry-After, then retry with backoff

This example supports the two standard forms of Retry-After: a number of seconds or an HTTP date. Without the header, it applies exponential backoff with a configurable starting delay and cap. Those fallback values are client policy choices, not PageCrawl-published intervals.

import email.utils
import os
import random
import time
from datetime import datetime, timezone

import requests

URL = "https://pagecrawl.io/api/your-endpoint"  # Replace from the API reference.
TOKEN = os.environ["PAGECRAWL_API_TOKEN"]
MAX_ATTEMPTS = 5
BACKOFF_BASE_SECONDS = 1.0  # Client-selected fallback, not a PageCrawl limit.
BACKOFF_CAP_SECONDS = 30.0  # Client-selected fallback.


def retry_after_seconds(value):
    if not value:
        return None
    try:
        return max(0.0, float(value))
    except ValueError:
        try:
            target = email.utils.parsedate_to_datetime(value)
            if target.tzinfo is None:
                target = target.replace(tzinfo=timezone.utc)
            return max(0.0, (target - datetime.now(timezone.utc)).total_seconds())
        except (TypeError, ValueError, OverflowError):
            return None


session = requests.Session()
headers = {"Authorization": f"Bearer {TOKEN}"}

for attempt in range(MAX_ATTEMPTS):
    response = session.get(URL, headers=headers, timeout=30)
    if response.status_code != 429:
        response.raise_for_status()
        print(response.text)
        break

    if attempt == MAX_ATTEMPTS - 1:
        response.raise_for_status()

    server_wait = retry_after_seconds(response.headers.get("Retry-After"))
    if server_wait is not None:
        wait = server_wait
    else:
        wait = min(BACKOFF_CAP_SECONDS, BACKOFF_BASE_SECONDS * (2 ** attempt))
        wait += random.uniform(0, min(1.0, wait * 0.1))
    time.sleep(wait)
else:
    raise RuntimeError("Retry loop ended unexpectedly")

Use this pattern only for operations safe to repeat. For a write or push operation, confirm the endpoint’s idempotency behavior in its documentation before automatically retrying: a network failure can leave the outcome uncertain even when the client did not receive the response.

Node.js: parse Retry-After and back off

This Node.js example uses built-in fetch. Set the endpoint from the API reference. The fallback delay values are example client settings; they are not PageCrawl limits.

const endpoint = 'https://pagecrawl.io/api/your-endpoint'; // Replace from the API reference.
const token = process.env.PAGECRAWL_API_TOKEN;
const maxAttempts = 5;
const backoffBaseMs = 1000; // Client-selected fallback.
const backoffCapMs = 30000; // Client-selected fallback.

function retryAfterMs(value) {
  if (!value) return null;
  const seconds = Number(value);
  if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
  const dateMs = Date.parse(value);
  if (!Number.isNaN(dateMs)) return Math.max(0, dateMs - Date.now());
  return null;
}

for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
  const response = await fetch(endpoint, {
    headers: { Authorization: `Bearer ${token}` },
  });

  if (response.status !== 429) {
    if (!response.ok) {
      throw new Error(`PageCrawl returned HTTP ${response.status}: ${await response.text()}`);
    }
    console.log(await response.text());
    break;
  }

  if (attempt === maxAttempts - 1) {
    throw new Error(`PageCrawl still returned HTTP 429 after ${maxAttempts} attempts`);
  }

  const serverWait = retryAfterMs(response.headers.get('retry-after'));
  let waitMs;
  if (serverWait !== null) {
    waitMs = serverWait;
  } else {
    const exponential = Math.min(backoffCapMs, backoffBaseMs * (2 ** attempt));
    waitMs = exponential + Math.random() * Math.min(1000, exponential * 0.1);
  }
  await new Promise((resolve) => setTimeout(resolve, waitMs));
}

In browser JavaScript, cross-origin access to response headers can be subject to CORS exposure rules. If your browser code cannot read Retry-After, make the API call from a server-side client or verify the API’s CORS behavior.

4. Prevent repeated rate limiting

Pace and queue work

When processing many URLs or records, put work in a queue and limit how quickly workers start requests. Use the current per-endpoint cap as an upper bound and leave room for other processes sharing the same account. A fixed per-minute allowance should not be treated as permission to send all requests at once: burst behavior may still trigger rate limiting.

Remove duplicate requests

Look for repeated polling, overlapping scheduled jobs, client retries layered on top of library retries, and duplicate submissions. PageCrawl says accepted Push API data pushes count toward the plan’s check allowance even when the value has not changed, although unchanged pushes deduplicate history entries. Avoiding redundant pushes can reduce unnecessary request pressure and allowance use.

Use webhooks when event delivery fits

For change notifications, compare polling through REST with receiving events through webhooks. REST polling is straightforward but each poll is a client request and your client owns retry handling. Webhooks can provide event-driven delivery without repeatedly polling; PageCrawl says webhook delivery automatically retries temporary failures with backoff. That webhook delivery behavior is separate from the rate limits on your own API requests.

5. Distinguish 429 from other errors

Status What PageCrawl documents What to do
429 Rate limited; Push API instructions say to honor Retry-After before retrying. Pause, follow the header if present, and reduce request pressure.
401 Invalid or missing API token on authenticated Push API endpoints. Check the Bearer token and authentication header. Do not retry rapidly with the same invalid credentials.
422 Validation error on the Push API. Inspect the response details and correct the request before retrying.

Status meanings here are from PageCrawl’s Push API documentation. Confirm the current endpoint reference for other API routes.

6. Troubleshooting common 429 problems

Symptom Likely cause Fix
429 appears after a short burst Concurrent workers or a batch exceeded the endpoint’s effective request rate. Queue requests and pace workers. Verify the cap for the exact endpoint and account.
429 repeats immediately after retry The client retries before Retry-After expires, or keeps producing requests at the same pace. Honor the header and slow or pause the producer while retries are pending.
No Retry-After header is visible The response may omit it, or an intermediary/client may not expose it. Capture raw response headers with a server-side request such as cURL. If absent, use spaced backoff rather than a tight loop.
Your plan appears to allow more requests than you can send The guides describe limits differently, or this endpoint has an account-specific limit. Check the current API reference and confirm with PageCrawl for the endpoint/account. Do not infer that a higher limit can be requested.
The error is actually 401 Missing or invalid Bearer token on an authenticated Push API route. Correct authentication before retrying.
The error is actually 422 The submitted data failed validation. Read the response details and fix the request fields.
Allowance is being used despite unchanged data Accepted Push API submissions count toward the plan’s check allowance even if the value did not change. Reduce duplicate pushes; unchanged values deduplicate history entries, but accepted pushes still count toward the allowance.

7. Performance, reliability, and cost considerations

  • Throughput: Controlled pacing lowers burst errors, but naturally bounds how quickly a backlog drains. Measure queue depth and completion time in your own system rather than assuming a universal request rate.
  • Reliability: Bounded retries prevent a failing request from occupying a worker indefinitely. Record status, endpoint, attempt count, and Retry-After so recurring throttling can be traced to the responsible workload.
  • Retry safety: A retry after a timeout can duplicate a write if the server processed the first request but its response was lost. Check the endpoint’s idempotency guarantees before retrying writes.
  • Allowance and cost: PageCrawl says REST API and webhook access are available on every plan, while request rates vary by plan in its API guide. For Push API data, accepted pushes count toward the plan’s check allowance even when the value is unchanged. Check your plan details and avoid unnecessary calls.

8. Or skip the browser setup

If your workflow also needs website screenshots, [ScreenshotNeo](https://screenshotneo.com) is a screenshot API and MCP server for developers. One GET request captures a URL as PNG, JPEG, WebP, or PDF. Its [API documentation](https://screenshotneo.com/docs/) covers the available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, and failed loads are never billed, and cache hits cost nothing. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server lets AI agents, including Claude, Cursor, and other MCP clients, take screenshots.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.

9. FAQ

Does HTTP 429 mean PageCrawl is down?

It means the request was rate-limited. Check the response and endpoint before diagnosing an outage.

How long should I wait?

Use the Retry-After value when it is present. PageCrawl does not publish one universal wait interval for every endpoint and account.

Can I request a higher rate limit?

The documentation reviewed does not establish whether increases are available or describe a request process. Ask PageCrawl to confirm your endpoint and account limit.

Should I switch from REST to webhooks?

Use webhooks when event delivery fits your workflow and you want to avoid repeated polling. They require handling incoming delivery and are distinct from client-initiated REST requests.

Sources