The Web Is Growing an Identity Layer for Bots: Signed Agents and What They Mean for Crawling
Signed-agent protocols let sites verify which bot key made an HTTP request. Learn what Web Bot Auth proves, how to deploy it, and where it falls short.
How can I verify that a bot visiting my website is legitimate? Use a signed-agent scheme based on HTTP Message Signatures, then make an independent access decision. A valid signature gives you evidence that a published agent identity controlled the signing key. It does not prove that the bot is harmless, authorized to read a page, or acting for a particular person.
Web Bot Auth is the emerging IETF approach for this identity layer. An automated client signs selected HTTP request components with a private key. The site discovers the corresponding public key through the agent’s published HTTPS identity and verifies the signature. The work is still an Internet-Draft, so implementations and header details can change.
What a signed agent proves
The protocol answers an identity question: which key, associated with which published agent identity, signed this request? That is stronger than trusting a self-reported User-Agent string or a source IP address alone.
- Identity evidence: the signature validates against a public key found through the agent’s key directory.
- Request integrity: the covered HTTP components have not changed since signing.
- Freshness: signature metadata can impose a bounded validity window.
A successful verification does not answer these separate policy questions:
- May this agent crawl this URL?
- Should it be rate-limited?
- Is the operator accountable to your terms?
- Is the request acting for a specific user?
Keep authorization, robots policy, rate limits, abuse controls and application permissions separate from signature verification. The OpenID Foundation’s agent-identity report also distinguishes public-web identity from workload identity used to authorize a permissioned API.
How Web Bot Auth fits together
- The agent generates or controls a private signing key.
- It publishes the matching public key in a JWKS-style directory.
- The request carries HTTP signature metadata and a Web Bot Auth agent identifier.
- Your origin, CDN or WAF discovers the key and verifies the covered request components.
- Your policy engine decides whether to allow, challenge, throttle, log or deny the request.
The current draft defines Signature-Agent as a way to discover candidate key material. The identifier is based on an HTTPS URL where the agent publishes keys. Signature metadata identifies the key and its validity window. Because the signature is carried in HTTP headers, deployment does not require changing TLS.
Read the current IETF Web Bot Auth draft when implementing a verifier. Earlier architecture drafts are historical context; do not use them as the wire-format authority.
What site operators should verify
1. Discover keys securely
Fetch the published directory over HTTPS, validate the response, and cache it according to the agent’s stated cache policy. Remove keys that disappear from the directory. Keep enough metadata to explain which key and directory version produced each decision.
2. Validate every required signature field
Verify both Signature and Signature-Input. Check the key identifier, algorithm, covered components, creation and expiration parameters, and the Web Bot Auth tag required by the draft. Reject malformed metadata instead of silently treating it as an ordinary unsigned request.
3. Confirm the request scope
Inspect exactly what was signed. A signature that covers only the authority may leave the method, path, query, headers or body outside the integrity boundary. For a crawler, method, target URI and any body or content-negotiation headers relevant to your policy should be covered according to the current draft.
4. Enforce freshness and replay controls
Use short validity windows and reject expired signatures. Do not treat a precomputed signature as a long-lived credential. If your application accepts requests with bodies or state-changing methods, add replay protections appropriate to that endpoint.
5. Make an explicit policy decision
Map verified identities to rules such as allowed paths, crawl budgets, authentication requirements and logging levels. An unknown key, an invalid signature and an unsigned request are different states and should be observable separately.
Origin, CDN or WAF verification
| Placement | Advantages | Trade-offs |
|---|---|---|
| Origin | Full application context, direct control over authorization and logs | Consumes origin capacity; every service must implement consistent verification |
| CDN or WAF | Checks traffic before it reaches the origin; shared key caching and rate controls | Requires provider support; application teams may see less context |
| Fronting proxy | Central policy and a gradual rollout across multiple origins | Needs careful forwarding of verified identity and decision headers |
Cloudflare documents Web Bot Auth as a cryptographic bot-authentication method built on IETF drafts. Provider behavior and supported fields can change, so confirm current documentation before making a vendor-specific configuration a requirement.
A practical verification flow
The following small Node.js service demonstrates the policy shape. It is intentionally limited to parsing and classifying metadata; production verification must implement the current draft’s HTTP Message Signature and key-discovery rules.
import http from 'node:http';
function classify(req) {
const agent = req.headers['signature-agent'];
const signature = req.headers['signature'];
const input = req.headers['signature-input'];
if (!agent || !signature || !input) {
return { state: 'unsigned', action: 'apply-fallback-policy' };
}
return {
state: 'needs-cryptographic-verification',
agent,
signaturePresent: true,
inputPresent: true
};
}
http.createServer((req, res) => {
const result = classify(req);
console.log(JSON.stringify({ method: req.method, url: req.url, result }));
res.writeHead(200, { 'content-type': 'application/json' });
res.end(JSON.stringify(result));
}).listen(8080, () => console.log('Listening on http://localhost:8080'));
Use a standards-compliant HTTP Message Signatures library once the draft version you target is fixed. Pin that version in deployment documentation, because the protocol is not final.
Testing with cURL
Use cURL to inspect how your edge handles signed and unsigned requests. Replace the URL with a test endpoint you control:
curl -i https://example.com/robots.txt
curl -i \
-H 'Signature-Agent: https://agent.example' \
-H 'Signature-Input: sig1=("@method" "@target-uri");created=1760000000;expires=1760000060' \
-H 'Signature: sig1=:BASE64_SIGNATURE:' \
https://example.com/robots.txt
The second command is a header-shape test, not a valid signature. Generate the value with your implementation and a private key; never copy a placeholder into production.
Unsigned and partially signed traffic
Do not block every request without a signature. Google’s experimental deployment signs only a subset of requests and says, We don’t sign every request of a particular agent.
Its guidance is to retain established verification methods and use IP-based verification as a fallback. Treat an absent signature as unknown, then combine available evidence:
- published IP ranges or reverse DNS, when available;
- robots.txt and documented crawl behavior;
- request rate, path patterns and error responses;
- application credentials for private APIs;
- historical reputation and abuse signals.
Key rotation and failure handling
- Unknown key: refresh the directory once, then mark the request unverifiable if the key is still absent.
- Expired key: reject verification and record the expiration reason.
- Withdrawn key: remove it from caches and do not continue accepting signatures made with it.
- Directory outage: use a bounded stale cache only if your risk policy permits it; otherwise fall back to unsigned-request controls.
- Clock skew: synchronize verifier clocks and allow only a documented, small tolerance.
- Malformed fields: return a clear verification result and avoid treating parser errors as proof of bad intent.
Privacy and correlation
A stable agent identity improves accountability and operational visibility. It can also enable cross-site correlation of the same crawler. Limit retained data, document who can access identity logs, and consider whether different activities need separate identities. Selective disclosure is discussed as a possible direction, but integrating it across the web remains difficult.
What the current evidence says
The MIT AI Agent Index examined a sample of 30 agents in 2025. Seven published stable User-Agent strings and IP ranges for verification, while six explicitly used Chrome-like User-Agent strings and residential or local IP contexts. These are sample findings, not a market-wide adoption rate.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Every request is unsigned | Provider signs only selected traffic | Keep IP and behavioral fallbacks; confirm provider coverage |
| Valid requests fail after rotation | Stale JWKS cache or removed key | Refresh the directory, evict withdrawn keys and honor cache policy |
| Signature verifies but path policy is wrong | Path or query was not covered | Require coverage of the target components used by authorization |
| Intermittent expiry failures | Clock skew or long network queues | Synchronize clocks and use a narrowly documented tolerance |
| Proxy sees different data than origin | Forwarded headers or URI were rewritten | Verify before rewriting, or define and sign the canonical representation |
| Bot is verified but abusive | Identity was mistaken for permission | Apply rate limits, robots rules and endpoint authorization separately |
Performance, reliability and cost
Key-directory fetches should be cached rather than performed for every request. Verification is CPU work, so perform it at a shared edge when traffic volume justifies it, while preserving the verified identity and reason code for the origin. Measure cache hit rate, verification latency, directory failures, expired signatures and fallback decisions. Avoid fail-open behavior for sensitive endpoints unless your threat model explicitly permits it.
The protocol itself does not define a universal pricing model. Your costs come from key-directory operations, signature verification, logging, edge requests and any bot-management service you add. Compare those costs with the operational burden of maintaining IP allowlists and pairwise shared secrets.
Or skip the browser setup
If your goal is to collect reliable page images while evaluating bot behavior, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL in one GET request and returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and each response reports the result in X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options. cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. It includes full-page capture, element selectors, device presets, custom headers and cookies, JavaScript, blocking rules, waiting conditions, caching, signed links, async webhooks, bulk capture and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account and start with 1,000 screenshots a month at no charge.
FAQ
Is Web Bot Auth a finished standard?
No. It is an IETF Internet-Draft and may change. Pin the draft version you implement and monitor revisions.
Does a valid signature mean I should allow the crawler?
No. It identifies the signer; your site still decides permission, rate limits and scope.
Should unsigned requests be blocked?
Usually not by default. Coverage can be partial, so combine signatures with existing verification and abuse controls.
Where should verification run?
At the origin when application context is essential, or at a CDN, WAF or fronting proxy when centralized filtering and shared key caching matter. Many deployments use both.
Can this replace API authentication?
No. Public-web agent identity and authorization to a private API are separate mechanisms. Continue using API credentials and endpoint authorization.
What is the main operational risk?
Incorrect key caching or an overly narrow signature scope. Both can reject valid traffic or give false confidence about what was protected.


