How to Troubleshoot API Errors
A practical, provider-neutral method to diagnose 400, 401, 403, 404, 429 and 5xx API errors, retry safely and escalate with useful evidence.

API errors become much easier to fix when you separate the failure into four questions: is the request valid, is the caller authenticated and authorized, has a limit been reached, and is the service temporarily unavailable? Start with the HTTP status, then read the provider’s response body, error code and headers. A status alone is only a clue: providers can use the same status for different causes. OpenAI’s error guide, GitHub’s REST guidance and Salesforce’s status documentation all show provider-specific differences.
Fast triage: identify the class of failure
| Status | Likely class | First checks |
|---|---|---|
| 400 | Malformed or invalid request | Method, endpoint version, parameters, headers, JSON syntax, required fields and types |
| 401 | Authentication | Credential present, active, unexpired, correctly scoped and sent in the required header or parameter |
| 403 | Permission or policy | Role and scopes, project or organization, IP restrictions, policy blocks and provider-specific limits |
| 404 | Missing or concealed resource | Path, identifier and API version; some services hide inaccessible resources as 404 |
| 409 | State conflict | Duplicate operation, stale version, resource state and idempotency rules |
| 422 | Well-formed but rejected data | Business validation, allowed values, relationships and field constraints |
| 429 | Rate, quota, credit or spending limit | Response code, Retry-After, rate headers, account limits and remaining credits |
| 500/502/503/504 | Server or upstream failure | Provider status, error detail, timeout behavior and safe retry policy |
This mapping is a triage aid, not a universal contract. Always follow the target API’s current documentation. Zoom specifically recommends inspecting the body’s provider error code and message alongside the HTTP status; Google Cloud documents its own authentication, resource and quota mappings.

1. Capture the complete failure
Before changing code, record one failed request in a sanitized form. Keep:
- HTTP method, full endpoint path and API version.
- Timestamp with time zone.
- Status code and complete response body.
- Non-secret response headers, especially request or correlation IDs, Retry-After and rate-limit fields.
- The exact input shape: query parameters, JSON field names and representative values.
- Whether the operation is read-only or can create, charge or mutate data.
Redact API keys, cookies, authorization headers, personal data and confidential payloads. The OpenAI Help Center gives the same escalation principle: do not include API keys or other authentication secrets.
# A safe diagnostic record (replace values with redacted examples)
request_id: req_abc123
method: POST
endpoint: https://api.example.com/v1/widgets
status: 429
time_utc: 2026-09-29T14:22:08Z
retry_after: 30
body_code: rate_limit_exceeded
2. Recheck the request contract (400 and 422)
A 400 usually means the server could not accept the request shape. Compare your request with the exact endpoint and version documentation:
- Confirm the HTTP method and URL path.
- Check required query and path parameters.
- Set the documented
Content-TypeandAcceptheaders. - Validate JSON syntax, nesting, capitalization and data types.
- Remove unsupported fields and verify enum values, date formats and size limits.
- Check that you are sending the body in the format the client library expects.
GitHub lists invalid JSON and invalid request structure as causes of 400 responses. A 422 often means the JSON is syntactically valid but violates a domain rule, such as a missing relationship or disallowed state.
curl -i -X POST 'https://api.example.com/v1/widgets' \
-H 'Authorization: Bearer REDACTED' \
-H 'Content-Type: application/json' \
--data '{"name":"example","enabled":true}'
Run the smallest valid request first, then add optional fields one at a time. This isolates the field that changes a success into a failure.
3. Diagnose 401, 403 and deceptive 404 responses
401 Unauthorized
Check that the credential is actually present in the process making the request, belongs to the intended account or project, has not expired or been revoked, and uses the required scheme such as Bearer. Verify environment selection: a local development key, staging key and production key may have different access.
403 Forbidden
A valid identity can still lack a scope, role or policy permission. Check resource-level access, organization membership, IP allowlists, regional restrictions and whether the endpoint requires an elevated plan. Some providers also use 403 for policy or abuse controls, so read the provider error code instead of assuming the token is wrong.
404 Not Found
First verify the path, identifier, API version and URL encoding. Then test access with the same identity. Some services intentionally return 404 when a private resource exists but the caller is not allowed to know that fact. GitHub documents this pattern. Do not conclude that the resource is absent until authentication and authorization have been checked.
4. Handle 429 without making the incident worse
A 429 can represent short-term request throttling, exhausted credits, a token quota or a spending limit. Read the body and headers before retrying. Retrying cannot restore an exhausted balance or a disabled spending limit.
If Retry-After is present and valid, wait at least that long. Otherwise use bounded exponential backoff with jitter. Cap both attempts and total elapsed time, and make sure your SDK is not already retrying underneath your application.
import random
import time
import requests
for attempt in range(5):
response = requests.get(
"https://api.example.com/v1/widgets",
headers={"Authorization": "Bearer REDACTED"},
timeout=30,
)
if response.status_code != 429:
response.raise_for_status()
data = response.json()
break
retry_after = response.headers.get("Retry-After")
if retry_after and retry_after.isdigit():
delay = float(retry_after)
else:
delay = min(60, 2 ** attempt) + random.uniform(0, 0.5)
time.sleep(delay)
else:
raise RuntimeError("Retry budget exhausted")
Use a queue or token bucket for sustained traffic. Spread scheduled jobs instead of releasing them all at once, and monitor remaining quota when the provider exposes it.
5. Treat 5xx and network failures as conditional retries
500, 502, 503 and 504 can be transient, but they can also expose a persistent server bug, invalid upstream dependency or timeout. Check the provider’s status page and returned error detail. OpenAI’s guidance distinguishes a brief wait for 500 from Retry-After-aware handling for 503 overload; that is provider guidance, not a universal contract.
Retry only operations that are safe to repeat. For writes, use the provider’s idempotency key or equivalent. A timeout does not prove that the server did nothing: the request may have succeeded while the response was lost. Before retrying a payment, job creation or deletion, query the operation status or use an idempotency mechanism.
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
async function getWithBackoff(url, options = {}) {
for (let attempt = 0; attempt < 4; attempt++) {
const res = await fetch(url, options);
if (![500, 502, 503, 504].includes(res.status)) return res;
const retryAfter = Number(res.headers.get('retry-after'));
const delay = Number.isFinite(retryAfter) && retryAfter >= 0
? retryAfter * 1000
: Math.min(30000, 500 * 2 ** attempt) + Math.random() * 250;
await sleep(delay);
}
throw new Error('Transient retry budget exhausted');
}
6. Isolate application, network and provider causes
Compare the failing application request with a minimal command-line request using the same endpoint and a redacted credential. If both fail, focus on the contract, identity, permissions, account limits or provider status. If the minimal request succeeds, inspect application serialization, environment variables, proxy and firewall rules, TLS certificates, DNS, redirects and retry middleware.

Log elapsed time and connection phase where your HTTP client supports it. A client-side timeout can occur before the provider receives a request, while a gateway timeout may mean an upstream operation is still running. Capture response headers on every failure; a request ID often lets support locate the server-side record.
7. Escalate with evidence
Send support the exact error text and provider code, request ID, occurrence time and time zone, applicable limit, sanitized endpoint and payload shape, and the troubleshooting steps already tried. Include whether retries changed the result. Never send API keys, cookies or authentication secrets.
Or skip the browser setup: troubleshoot screenshot API calls with ScreenshotNeo
If the API you are integrating is a website screenshot endpoint, ScreenshotNeo gives you a simple request and diagnostic headers that make failures explicit. It accepts a URL and returns PNG, JPEG, WebP or PDF. Clean shots are billed only after consent banners, newsletter popups and chat widgets are removed; bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing. Each response identifies the result with X-Page-Verdict and X-Billed headers.
Read the complete parameter list in the ScreenshotNeo API documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
print(r.headers.get("X-Page-Verdict"), r.headers.get("X-Billed"))
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', bytes);
console.log(res.headers.get('X-Page-Verdict'), res.headers.get('X-Billed'));
For screenshot-specific errors, first verify the access key, URL encoding and response status. Then inspect X-Page-Verdict and X-Billed: a bot check, blank page, timeout or failed load explains why an image was not produced and is not billed. If a page needs authentication, pass supported custom headers, cookies, user agent or Authorization settings. For dynamic pages, choose a wait for a selector, delay or network idle; for long pages, use full-page capture with lazy images loaded. You can also capture one element by CSS selector, apply custom CSS or JavaScript, click before capture, hide selectors, block selected requests or resource types, set timezone or geolocation, use a device preset or arbitrary viewport, and choose retina scale.
For reliability and cost, use a cache TTL when repeated URLs are acceptable, bulk capture for up to 100 URLs per call, and asynchronous jobs with signed webhooks for long-running batches. Only clean shots are billed. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Create a free ScreenshotNeo account with 1,000 screenshots per month and no card.
Performance, reliability and cost checklist
- Set a client timeout that covers expected server work, but keep a total retry deadline.
- Reuse HTTP connections and enable compression where supported.
- Use bounded concurrency and a queue instead of an unbounded worker pool.
- Honor Retry-After and add jitter to fallback delays.
- Make writes idempotent before enabling automatic retries.
- Record request IDs, status, provider codes and latency without logging secrets.
- Cache immutable reads and screenshot results when freshness permits.
- Separate provider errors from client validation errors in metrics and alerts.
FAQ
Should I retry every error?
No. Fix malformed requests and credentials first. Retry only documented transient failures, and make mutations idempotent.
Why did a valid resource return 404?
The path may be wrong, or the provider may conceal an inaccessible resource as 404. Test with the intended identity and verify scopes.
Does a 429 always mean I am sending requests too quickly?
No. It can indicate rate throttling, exhausted credits, quota or spending controls. Read the body and headers before retrying.
What should I give support?
Provide the sanitized request shape, exact error code and message, request ID, timestamp with time zone, relevant limit and steps already attempted. Omit secrets.
How can I tell whether a timeout changed data?
Check operation status or use the provider’s idempotency key. A lost response does not prove that the server rolled back the operation.


