What Is HTTP 403 in Web Scraping?
HTTP 403 means a server understood your scraping request but refuses it. Learn what it signals, how to diagnose it, and what to do next.

HTTP 403 Forbidden means the server understood your request but refuses to fulfill it. In web scraping, that status is a decision from the server, not a diagnosis of one universal problem. The response body and headers may explain the refusal, but a 403 does not automatically mean your password is wrong, that you are rate limited, or that changing one header will solve it.
RFC 9110 defines 403 this way: “The 403 (Forbidden) status code indicates that the server understood the request but refuses to fulfill it.” See RFC 9110, section 15.5.4. The practical response is to inspect the server’s explanation, confirm that your crawler is authorized, and use the site’s documented access route. If the refusal continues, stop automated requests rather than treating retries, proxy rotation, or browser imitation as guaranteed fixes.
What a 403 tells you
A 403 confirms that the server received and understood the request at the HTTP level. It then chose not to provide the requested representation. The reason can be related to account permissions, a crawler policy, a network rule, a security product, the requested method, or another site-specific decision.

Credentials are only one possibility. RFC 9110 explains that when credentials were supplied, the server may consider them insufficient, but the refusal can also be unrelated to credentials. A client should not automatically repeat the request with the same credentials.
Read the response as a signal:
- Status: record the exact code and request method.
- Body: look for an explanation, policy link, request ID, or instructions.
- Headers: preserve fields such as
Retry-After,WWW-Authenticate, server identifiers, and correlation IDs. - Request context: save the URL, account, timestamp, and authorization scope without logging secrets.
403 compared with nearby HTTP statuses
| Status | Meaning | What it suggests |
|---|---|---|
| 401 Unauthorized | Authentication credentials are absent or invalid. | The response normally includes a WWW-Authenticate challenge. Obtain or correct credentials through the documented flow. |
| 403 Forbidden | The server understood the request but refuses it. | Authorization, policy, crawler controls, request context, or another site-specific rule may be involved. The code alone cannot identify which. |
| 404 Not Found | The server did not find a current representation, or is unwilling to disclose that one exists. | Check the URL, routing, publication state, and account visibility. |
| 429 Too Many Requests | The client sent too many requests in a period of time. | Apply the server’s documented limits and honor any Retry-After value. |
| 503 Service Unavailable | The service is temporarily unable to handle the request, often due to overload or maintenance. | Use bounded retries only when appropriate and follow Retry-After if supplied. |
These categories can overlap operationally. A security gateway might return 403 for a request that another system would describe as throttled. Treat the body, headers, and published policy as the site-specific evidence.
Why web scrapers receive 403 responses
Account or resource permissions
Your account may not be entitled to the page, API route, tenant, geographic region, or HTTP method. Confirm the resource is intended for your account and that the token has the required scope. Do not assume that a valid login grants access to every URL.
Crawler policy or security controls
Sites may apply rules to automated clients, networks, request patterns, or protected routes. A security service can make its own decision before the application handles your request. The response body may identify the provider or provide a support reference.
Incorrect request details
An unexpected method, missing required parameter, wrong host, invalid path, expired signed URL, or malformed authorization value can produce a 403. Compare your request with the site’s official API documentation or a known-good example for your account.
Application-level denial
The application may deliberately return 403 when a user is not allowed to view an object. This can prevent disclosure of whether a private resource exists. A 403 therefore does not prove that the URL is public or that the content is present.
A responsible 403 troubleshooting sequence
- Capture evidence once. Save the status, response body, relevant headers, method, URL, and timestamp. Redact cookies, tokens, and personal data before sharing logs.
- Confirm the target. Check the scheme, hostname, path, query parameters, redirects, and HTTP method. Make sure the resource is available to the account or crawler you are using.
- Check authorization. Verify that credentials are current, sent in the documented format, and authorized for this resource. Do not keep replaying the same credentials after a refusal.
- Read the policy. Consult the site’s API terms, crawler instructions, and contact route. RFC 9309 defines robots.txt rules for automated clients, but explicitly says: “These rules are not a form of access authorization.” Following robots.txt does not grant permission, and its absence does not grant permission either.
- Reduce your request rate while investigating. Avoid bursts that create more load or obscure the original evidence. Keep a bounded diagnostic sample.
- Ask the owner or use an approved API. If the refusal persists or permission is unclear, stop automated collection and request access or use an authorized data source.
Inspecting a 403 with cURL
Use -i to display response headers and -L only when following redirects is appropriate. The command records the body for inspection; it does not attempt to bypass the refusal.
curl -i --max-time 30 'https://example.com/private-page' -o response.txt
# If you need headers and body in separate files:
curl -sS -D response.headers --max-time 30 \
'https://example.com/private-page' \
-o response.body
Review response.body for a policy message or request ID. Check response.headers for Retry-After, WWW-Authenticate, and provider-specific support fields.
Inspecting a 403 in Python
Python’s standard library names 403 FORBIDDEN in http.HTTPStatus. This example prints safe diagnostics and avoids exposing authorization headers.
from http import HTTPStatus
import requests
url = 'https://example.com/private-page'
response = requests.get(url, timeout=30)
print('status:', response.status_code, response.reason)
print('is_forbidden:', response.status_code == HTTPStatus.FORBIDDEN)
print('content_type:', response.headers.get('content-type'))
print('retry_after:', response.headers.get('retry-after'))
print('body_preview:', response.text[:500])
if response.status_code == HTTPStatus.FORBIDDEN:
print('The server refused this request; check authorization and policy.')
response.raise_for_status()
For an authenticated API, pass credentials exactly as documented, keep them in environment variables, and log only a redacted request description. A 403 should be handled as an authorization or policy decision, not as an invitation to cycle through headers.
Inspecting a 403 in Node.js
const url = 'https://example.com/private-page';
const response = await fetch(url, { signal: AbortSignal.timeout(30000) });
const body = await response.text();
console.log({
status: response.status,
contentType: response.headers.get('content-type'),
retryAfter: response.headers.get('retry-after'),
bodyPreview: body.slice(0, 500)
});
if (response.status === 403) {
throw new Error('The server refused this request; verify authorization and policy.');
}
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
Robots.txt, authorization, and compliance
Robots.txt is crawler guidance, not a permission document. RFC 9309 describes how automated clients can discover and follow requested rules, while making clear that the protocol does not authorize access. You still need permission, valid credentials where required, and compliance with the site’s terms and applicable law.
A commercial crawling or browser service can help operate an authorized workflow, but no service can grant permission to access a particular site. If you cannot establish authorization, use a public API, licensed dataset, or another approved source.
Common mistakes and their fixes
| Mistake | Why it fails | Better action |
|---|---|---|
| Assuming 403 always means a bad password | RFC 9110 allows refusals unrelated to credentials. | Read the body and verify the resource and policy. |
| Retrying unchanged credentials indefinitely | The server has already refused that authorization context. | Stop, correct the authorization path, or contact the owner. |
| Treating robots.txt as a green light | Robots rules are not access authorization. | Obtain permission separately. |
| Calling 403 rate limiting | 429 is the defined Too Many Requests status, although a gateway may choose another code. | Follow the actual response and published limits. |
| Changing user-agent or rotating proxies as a “fix” | Those changes do not establish authorization and may violate policy. | Use an approved API or request access. |
| Ignoring the response body | The server may provide the only useful explanation there. | Store and inspect the representation safely. |
Performance and reliability considerations
A reliable collector treats refusals as first-class outcomes. Record status counts, response sizes, latency, and request IDs. Use a queue with bounded concurrency, exponential backoff only for statuses and endpoints where retries are documented, and a dead-letter path for persistent 403 responses. Do not retry a permanent authorization denial on every scheduled run.
Cache successful responses where freshness permits. Separate transient 503 handling from 403 handling so an unavailable service does not become an endless permission loop. Keep a per-domain budget and stop when the site signals that your collection is not wanted. These controls reduce load and make an audit trail possible.
Or skip the browser setup
If your goal is a clean visual capture rather than raw HTML extraction, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP, or PDF. The direct request is:
curl -G 'https://api.screenshotneo.com/v1/shot' \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for parameters and response details. Consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing result. ScreenshotNeo also provides an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools.
The free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
FAQ
Is 403 the same as being rate limited?
No. 429 is the standard status for Too Many Requests. A particular gateway can use 403 for a policy decision, so inspect its body and documentation.

Can a browser solve every scraping 403?
No. Browser rendering may change the request context, but it does not create permission or guarantee access. Follow the site’s approved route.
Should I keep requesting after a 403?
Only when the owner’s documentation explicitly defines a safe, authorized retry path. Otherwise stop and resolve authorization or contact the site.
Does a 403 prove that the page exists?
No. A server may use 403 instead of revealing whether a protected resource exists.
What should I include in a support request?
Provide the timestamp, URL and method, status, redacted headers, response body excerpt, account or application identifier, and request ID. Never include secrets.