ScreenshotNeo

BlogGuides

HTTP Referer Header: A Complete Guide for Web Scraping

Learn what the HTTP Referer header tells a server, when to send it in scrapers, and how to handle privacy, redirects, errors, and browser captures.

By the ScreenshotNeo team29 September 20269 min read

HTTP Referer Header: A Complete Guide for Web Scraping

The HTTP Referer request header optionally identifies the URI from which a request’s target was obtained. For scraping, treat it as optional request metadata: send it only when it accurately describes the request context or the destination’s documented requirements. It is not proof that a human visited a page, an access credential, or permission to scrape. The field name preserves a historical misspelling; “referrer” is the ordinary spelling and appears in Referrer-Policy. RFC 9110 §10.1.3

This guide explains how to set, omit, and inspect the header in Python, Node.js, and cURL; how browser referrer policies affect it; and what to do when a site behaves differently than expected.

1. What the Referer header contains

When a user agent sends the field, its value is a URI reference for the resource from which the target URI was obtained. The value can be an absolute URI or a partial URI. A conforming user agent generating it leaves out the URI fragment and userinfo components. For example, navigation from https://example.com/catalog/item?color=blue#details could produce a value that omits #details; policy and client behavior may reduce it further. RFC 9110

The Referer field can describe a request’s source, but it is optional metadata rather than proof of access.
The Referer field can describe a request’s source, but it is optional metadata rather than proof of access.

Servers may use the field for backlink generation, analytics, link maintenance, caching decisions, or request checks. Its presence is not guaranteed: not all requests include it, and a user agent may truncate information beyond the referring origin. A missing value does not prove there was no referring page. A present value is request metadata, not reliable evidence of identity or a genuine browser journey. RFC 9110 §10.1.3

2. Should a scraper send Referer?

Usually, omit it unless you have a real referring page for the request or the destination documents a legitimate requirement. Do not invent a value just to make an automated request look like a human navigation. Fabricated provenance can be misleading, and the header grants no rights.

Situation Practical choice
Your client fetches a known link from a page you actually retrieved You may send that page’s URI if doing so reflects the request flow and privacy policy.
You make a direct request with no referring resource Omit the header unless the service documents a specific requirement.
A site returns an error mentioning referrer validation Check its official API or access documentation; do not guess at a bypass.
You are auditing or analyzing referrer behavior Record whether the header was sent, its value, redirect hops, and response status.
You are crawling pages Review the site’s terms and crawler guidance separately. A Referer value and an allowed robots.txt path do not authorize access.

Robots rules are requested crawler behavior, not access authorization. Respect the destination’s published policies and applicable access controls. RFC 9309 §1

3. Send or omit Referer with Python

Install Requests with python -m pip install requests. This runnable example sends a truthful referring URI, prints the response status and returned header, and has a timeout. Replace the two example URLs with resources you are permitted to access.

import requests

source_page = "https://example.com/catalog"
target_url = "https://example.com/catalog/item"

response = requests.get(
    target_url,
    headers={"Referer": source_page},
    timeout=(5, 30),
)
print("status:", response.status_code)
print("sent referer:", response.request.headers.get("Referer"))
print("content type:", response.headers.get("Content-Type"))
response.raise_for_status()
print(response.text[:500])

To omit it, remove the headers argument. If a session is needed for cookies across requests, scope the header to the specific request rather than setting a global default without a reason:

with requests.Session() as session:
    first = session.get("https://example.com/catalog", timeout=(5, 30))
    first.raise_for_status()

    target = session.get(
        "https://example.com/catalog/item",
        headers={"Referer": first.url},
        timeout=(5, 30),
    )
    target.raise_for_status()
    print(target.status_code, target.url)

When following redirects, inspect response.url and response.history. Do not assume that every HTTP library applies browser navigation rules or changes a manually supplied header in exactly the same way. If redirect behavior matters, test the client you deploy and log each hop without exposing sensitive query strings.

4. Send or omit Referer with Node.js

With a current Node.js runtime that provides the global fetch, pass a header in the request options. This example has an abort timeout and checks for HTTP errors:

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 30_000);

try {
  const response = await fetch("https://example.com/catalog/item", {
    headers: { Referer: "https://example.com/catalog" },
    signal: controller.signal,
    redirect: "follow",
  });

  console.log("status:", response.status);
  console.log("final URL:", response.url);
  console.log("content type:", response.headers.get("content-type"));

  if (!response.ok) {
    throw new Error(`HTTP ${response.status}`);
  }
  const html = await response.text();
  console.log(html.slice(0, 500));
} finally {
  clearTimeout(timer);
}

To omit the field, remove headers. For a redirect audit, use redirect: "manual" and inspect the Location response header yourself, following only destinations your crawler is meant to access. Runtime implementations can differ from browsers, so validate the actual request headers when precise behavior matters.

5. Set Referer with cURL

Use -H to provide a value explicitly. This command prints response headers and saves the body for inspection:

curl --verbose \
  -H 'Referer: https://example.com/catalog' \
  --max-time 30 \
  --dump-header response-headers.txt \
  --output page.html \
  'https://example.com/catalog/item'

To omit the field, remove the -H option. cURL may have other defaults or configuration files, so verbose output helps confirm what was sent. Avoid sharing verbose logs without reviewing them for cookies, authorization values, and private URLs.

6. Browser privacy and Referrer-Policy

A page can control referrer disclosure through the Referrer-Policy response header, a meta element, supported element-level referrerpolicy attributes, or noreferrer. Policies include no-referrer, same-origin, origin, strict-origin, origin-when-cross-origin, strict-origin-when-cross-origin, no-referrer-when-downgrade, and unsafe-url. These choices control whether the header is omitted or how much URI information is disclosed. W3C Referrer Policy

Referrer policies control how much source information a browser discloses on a request.
Referrer policies control how much source information a browser discloses on a request.
Policy example General effect
no-referrer Do not send referrer information.
same-origin Send it for same-origin requests only.
origin Send the origin rather than the full path.
strict-origin-when-cross-origin Typically sends full information same-origin and origin information cross-origin, subject to secure-to-insecure downgrade restrictions.
unsafe-url Can disclose the full referring URL, including path and query, so it deserves careful privacy review.

The exact behavior depends on the browser context, policy, and request. The W3C specification describes no-referrer-when-downgrade as a default in the behavior covered by that report; do not treat this as an evergreen guarantee for every current browser. Check current browser and Fetch documentation for a specific deployment.

RFC 9110 says a user agent must not send Referer in an unsecured HTTP request when the referring resource was accessed securely. For secure cross-origin requests, it says the user agent should not send it unless the source explicitly allows it. The field can expose personal or confidential information in a path or query. RFC 9110 §10.1.3

7. Scraping workflow and edge cases

  1. Establish the request context. Determine whether the target came from a page you fetched, a user-provided URL, or a direct API call.
  2. Choose the minimum disclosure. Omit the field when no truthful referring resource exists. If one does, consider whether the full URL is necessary or whether the origin is enough.
  3. Keep request state explicit. Record the method, target, header policy, cookies, redirect chain, and response status. Do not log secret-bearing query parameters indiscriminately.
  4. Handle redirects deliberately. A redirect changes the target. Review the final destination and client behavior; avoid forwarding credentials or sensitive context to unrelated origins.
  5. Separate identity and access controls. Referer is not authentication. Use documented API credentials where authorized, and do not treat robots.txt as permission.
  6. Re-check when behavior changes. A server-side change, a browser policy, or a client library update can alter whether or how this metadata is sent.

For a multi-step crawl, each request may have a different genuine source. Avoid assigning one fixed referrer to every URL just to satisfy a server check. If a destination requires a particular workflow, use its documented interface and access process.

8. Troubleshooting common problems

Symptom Likely cause What to do
Server logs show no Referer The client omitted it, policy suppressed it, or the request has no referring context. Inspect the outgoing request. Add a value only when it accurately represents the source or documented service requirements.
Only the domain appears A privacy policy or client behavior reduced the value to an origin. Check the source page’s policy and the request context. Do not assume the full path should be available.
HTTPS-to-HTTP request lacks the header Secure-to-insecure disclosure is restricted for user agents. Keep requests on HTTPS where possible; do not weaken privacy expectations to force disclosure.
Header differs after a redirect Redirect handling and cross-origin policy affect subsequent requests. Inspect the redirect chain and test the exact HTTP client. Avoid carrying private URL data across origins.
403 or 404 persists after setting Referer The cause may be authentication, authorization, an invalid path, rate limiting, or another server rule. Read the service documentation and response body. A Referer value does not grant access; use supported access methods.
Scraper output leaks query data A full referring URL may contain identifiers or confidential query parameters. Omit the field or use the least revealing truthful value. Redact logs and review retention.
Python or Node sends unexpected headers Library defaults, session state, proxies, or configuration may affect the request. Inspect the prepared request or server-side capture in a controlled environment; keep the deployed runtime and dependency versions pinned and reviewed.

9. Performance, reliability, and cost

The Referer header is a small piece of request metadata; setting it does not make a scraper faster, more reliable, or more browser-like by itself. Request latency and reliability depend on the destination, network, timeouts, retries, rendering needs, and response handling. Use bounded timeouts, retry transient failures with backoff where appropriate, and avoid retrying permanent authorization or policy errors blindly.

For large crawls, keep concurrency within the destination’s published limits, honor rate guidance, cache responses when appropriate, and collect status and failure data. Do not infer that a successful HTTP response contains the intended page: verify status, content type, and expected content. A request header does not solve JavaScript rendering, consent banners, or dynamic page state.

10. Browser screenshots instead of raw HTTP

Use a browser-based capture when the task needs the rendered page rather than its HTML response, such as a visual archive, QA record, or image for documentation. A screenshot still does not bypass a site’s access controls or policies. Referer policy remains a browser disclosure mechanism; do not add false navigation context.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. One GET request returns a PNG, JPEG, WebP, or PDF capture. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners are accepted and removed before the capture, along with supported newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, no card required.

11. Short FAQ

Why is it spelled Referer?

The HTTP field name retains a historical misspelling. The policy mechanism uses the conventional spelling, as in Referrer-Policy.

Can I use Referer to prove a request came from my site?

No. It is optional metadata and can be omitted or constrained. Do not use it as sole proof of identity or authorization.

Does robots.txt tell me whether I am authorized?

No. RFC 9309 explicitly says robots rules are not access authorization. Follow site terms and access controls independently.

Should every scraper imitate a browser’s headers?

No universal header set is established here. Use the destination’s documented requirements and send only accurate, necessary metadata.

Primary references