PageCrawl.io API rate limits: how to handle 429 errors
A 429 means PageCrawl is rate-limiting your requests. Learn how to honor Retry-After, add backoff, pace calls, and check the limit for your endpoint and account.
HTTP 429 means PageCrawl.io is rate-limiting requests to an endpoint. Stop sending requests at the same pace, honor the Retry-After response header when it is present, and use backoff for automated retries. Do not retry immediately in a tight loop.
The applicable request ceiling depends on the endpoint and account. PageCrawl’s developer guide says its REST API allows 60 requests per minute on Free and 300 per minute on paid plans, while a separate dashboard guide describes 60 requests per minute as typical for most accounts. Check the current [PageCrawl API reference](https://pagecrawl.io/developers) for the endpoint and account you are using rather than treating either figure as universal.
1. What a 429 response tells you
A 429 response indicates that the server is limiting the request rate. By itself, it does not establish that the API is down. The Push API documentation instructs clients to honor Retry-After before retrying: “429 | Rate limited; honor the Retry-After header before retrying”.
PageCrawl documents different request limits in different guides. Its API developer guide lists 60 requests per minute for Free and 300 per minute for paid plans; its dashboard guide says most accounts have a 60-requests-per-minute limit. These are vendor-documented figures, not independent measurements. The API reference is described by PageCrawl as generated from its OpenAPI specification and taking precedence over guide text.
2. Respond to 429 in order
- Confirm the status. Record the request URL or endpoint, timestamp, account and plan context, response body, and response headers. Make sure the response is actually 429.
- Honor
Retry-After. If supplied, wait as directed before retrying. Do not substitute a guessed fixed delay for the server’s instruction. - Stop the current request burst. Pause or slow the producer of requests so new work does not immediately create another burst while retries are pending.
- Retry with backoff. If no
Retry-Afterheader is present, use a backoff policy that spaces retries and avoids a synchronized tight loop. PageCrawl’s dashboard guide recommends exponential backoff but does not establish a universal delay to use. - Reduce unnecessary calls. Check for duplicate work and pace requests below the verified endpoint/account limit.
- Verify the applicable limit. Consult the current API reference for the endpoint and your account. If the published information does not resolve your situation, ask PageCrawl for confirmation; the reviewed documentation does not establish a process for requesting a higher limit.
3. Runnable retry examples
The examples below demonstrate the retry pattern for an authenticated PageCrawl request. Replace the illustrative endpoint and payload with the endpoint and request format from the current API reference. Do not assume every endpoint uses the same method or payload. PageCrawl documents Bearer token authentication for its authenticated Push API endpoints.
cURL: inspect a response and its headers
curl -i \
-H "Authorization: Bearer $PAGECRAWL_API_TOKEN" \
"https://pagecrawl.io/api/your-endpoint"
Use the actual endpoint shown in the API reference. The -i option includes response headers, so you can inspect the status and Retry-After. cURL does not implement a general policy here; check the header and wait before issuing another request.
Python: wait for Retry-After, then retry with backoff
This example supports the two standard forms of Retry-After: a number of seconds or an HTTP date. Without the header, it applies exponential backoff with a configurable starting delay and cap. Those fallback values are client policy choices, not PageCrawl-published intervals.
import email.utils
import os
import random
import time
from datetime import datetime, timezone
import requests
URL = "https://pagecrawl.io/api/your-endpoint" # Replace from the API reference.
TOKEN = os.environ["PAGECRAWL_API_TOKEN"]
MAX_ATTEMPTS = 5
BACKOFF_BASE_SECONDS = 1.0 # Client-selected fallback, not a PageCrawl limit.
BACKOFF_CAP_SECONDS = 30.0 # Client-selected fallback.
def retry_after_seconds(value):
if not value:
return None
try:
return max(0.0, float(value))
except ValueError:
try:
target = email.utils.parsedate_to_datetime(value)
if target.tzinfo is None:
target = target.replace(tzinfo=timezone.utc)
return max(0.0, (target - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError, OverflowError):
return None
session = requests.Session()
headers = {"Authorization": f"Bearer {TOKEN}"}
for attempt in range(MAX_ATTEMPTS):
response = session.get(URL, headers=headers, timeout=30)
if response.status_code != 429:
response.raise_for_status()
print(response.text)
break
if attempt == MAX_ATTEMPTS - 1:
response.raise_for_status()
server_wait = retry_after_seconds(response.headers.get("Retry-After"))
if server_wait is not None:
wait = server_wait
else:
wait = min(BACKOFF_CAP_SECONDS, BACKOFF_BASE_SECONDS * (2 ** attempt))
wait += random.uniform(0, min(1.0, wait * 0.1))
time.sleep(wait)
else:
raise RuntimeError("Retry loop ended unexpectedly")
Use this pattern only for operations safe to repeat. For a write or push operation, confirm the endpoint’s idempotency behavior in its documentation before automatically retrying: a network failure can leave the outcome uncertain even when the client did not receive the response.
Node.js: parse Retry-After and back off
This Node.js example uses built-in fetch. Set the endpoint from the API reference. The fallback delay values are example client settings; they are not PageCrawl limits.
const endpoint = 'https://pagecrawl.io/api/your-endpoint'; // Replace from the API reference.
const token = process.env.PAGECRAWL_API_TOKEN;
const maxAttempts = 5;
const backoffBaseMs = 1000; // Client-selected fallback.
const backoffCapMs = 30000; // Client-selected fallback.
function retryAfterMs(value) {
if (!value) return null;
const seconds = Number(value);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1000);
const dateMs = Date.parse(value);
if (!Number.isNaN(dateMs)) return Math.max(0, dateMs - Date.now());
return null;
}
for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
const response = await fetch(endpoint, {
headers: { Authorization: `Bearer ${token}` },
});
if (response.status !== 429) {
if (!response.ok) {
throw new Error(`PageCrawl returned HTTP ${response.status}: ${await response.text()}`);
}
console.log(await response.text());
break;
}
if (attempt === maxAttempts - 1) {
throw new Error(`PageCrawl still returned HTTP 429 after ${maxAttempts} attempts`);
}
const serverWait = retryAfterMs(response.headers.get('retry-after'));
let waitMs;
if (serverWait !== null) {
waitMs = serverWait;
} else {
const exponential = Math.min(backoffCapMs, backoffBaseMs * (2 ** attempt));
waitMs = exponential + Math.random() * Math.min(1000, exponential * 0.1);
}
await new Promise((resolve) => setTimeout(resolve, waitMs));
}
In browser JavaScript, cross-origin access to response headers can be subject to CORS exposure rules. If your browser code cannot read Retry-After, make the API call from a server-side client or verify the API’s CORS behavior.
4. Prevent repeated rate limiting
Pace and queue work
When processing many URLs or records, put work in a queue and limit how quickly workers start requests. Use the current per-endpoint cap as an upper bound and leave room for other processes sharing the same account. A fixed per-minute allowance should not be treated as permission to send all requests at once: burst behavior may still trigger rate limiting.
Remove duplicate requests
Look for repeated polling, overlapping scheduled jobs, client retries layered on top of library retries, and duplicate submissions. PageCrawl says accepted Push API data pushes count toward the plan’s check allowance even when the value has not changed, although unchanged pushes deduplicate history entries. Avoiding redundant pushes can reduce unnecessary request pressure and allowance use.
Use webhooks when event delivery fits
For change notifications, compare polling through REST with receiving events through webhooks. REST polling is straightforward but each poll is a client request and your client owns retry handling. Webhooks can provide event-driven delivery without repeatedly polling; PageCrawl says webhook delivery automatically retries temporary failures with backoff. That webhook delivery behavior is separate from the rate limits on your own API requests.
5. Distinguish 429 from other errors
| Status | What PageCrawl documents | What to do |
|---|---|---|
| 429 | Rate limited; Push API instructions say to honor Retry-After before retrying. |
Pause, follow the header if present, and reduce request pressure. |
| 401 | Invalid or missing API token on authenticated Push API endpoints. | Check the Bearer token and authentication header. Do not retry rapidly with the same invalid credentials. |
| 422 | Validation error on the Push API. | Inspect the response details and correct the request before retrying. |
Status meanings here are from PageCrawl’s Push API documentation. Confirm the current endpoint reference for other API routes.
6. Troubleshooting common 429 problems
| Symptom | Likely cause | Fix |
|---|---|---|
| 429 appears after a short burst | Concurrent workers or a batch exceeded the endpoint’s effective request rate. | Queue requests and pace workers. Verify the cap for the exact endpoint and account. |
| 429 repeats immediately after retry | The client retries before Retry-After expires, or keeps producing requests at the same pace. |
Honor the header and slow or pause the producer while retries are pending. |
No Retry-After header is visible |
The response may omit it, or an intermediary/client may not expose it. | Capture raw response headers with a server-side request such as cURL. If absent, use spaced backoff rather than a tight loop. |
| Your plan appears to allow more requests than you can send | The guides describe limits differently, or this endpoint has an account-specific limit. | Check the current API reference and confirm with PageCrawl for the endpoint/account. Do not infer that a higher limit can be requested. |
| The error is actually 401 | Missing or invalid Bearer token on an authenticated Push API route. | Correct authentication before retrying. |
| The error is actually 422 | The submitted data failed validation. | Read the response details and fix the request fields. |
| Allowance is being used despite unchanged data | Accepted Push API submissions count toward the plan’s check allowance even if the value did not change. | Reduce duplicate pushes; unchanged values deduplicate history entries, but accepted pushes still count toward the allowance. |
7. Performance, reliability, and cost considerations
- Throughput: Controlled pacing lowers burst errors, but naturally bounds how quickly a backlog drains. Measure queue depth and completion time in your own system rather than assuming a universal request rate.
- Reliability: Bounded retries prevent a failing request from occupying a worker indefinitely. Record status, endpoint, attempt count, and
Retry-Afterso recurring throttling can be traced to the responsible workload. - Retry safety: A retry after a timeout can duplicate a write if the server processed the first request but its response was lost. Check the endpoint’s idempotency guarantees before retrying writes.
- Allowance and cost: PageCrawl says REST API and webhook access are available on every plan, while request rates vary by plan in its API guide. For Push API data, accepted pushes count toward the plan’s check allowance even when the value is unchanged. Check your plan details and avoid unnecessary calls.
8. Or skip the browser setup
If your workflow also needs website screenshots, [ScreenshotNeo](https://screenshotneo.com) is a screenshot API and MCP server for developers. One GET request captures a URL as PNG, JPEG, WebP, or PDF. Its [API documentation](https://screenshotneo.com/docs/) covers the available parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, timeouts, and failed loads are never billed, and cache hits cost nothing. Responses include
X-Page-VerdictandX-Billedheaders. - An MCP server lets AI agents, including Claude, Cursor, and other MCP clients, take screenshots.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.
9. FAQ
Does HTTP 429 mean PageCrawl is down?
It means the request was rate-limited. Check the response and endpoint before diagnosing an outage.
How long should I wait?
Use the Retry-After value when it is present. PageCrawl does not publish one universal wait interval for every endpoint and account.
Can I request a higher rate limit?
The documentation reviewed does not establish whether increases are available or describe a request process. Ask PageCrawl to confirm your endpoint and account limit.
Should I switch from REST to webhooks?
Use webhooks when event delivery fits your workflow and you want to avoid repeated polling. They require handling incoming delivery and are distinct from client-initiated REST requests.


