Identity and Fingerprinting for Browser Agents
Learn how websites identify browser agents, what fingerprinting reveals, how detection works, and how to reduce unnecessary exposure without breaking compatibility.
Direct answer: Websites identify browser agents by combining signals from HTTP requests, browser APIs, rendering behavior, network connections, and interaction patterns. A browser agent is software acting for a person. Fingerprinting is the process of combining these signals to distinguish or track a client. Bot detection is the service-side decision about whether that client appears automated, suspicious, or allowed. These are related but different ideas.
No single field proves that a visitor is an AI agent. A User-Agent string can be changed, and a browser context can be configured with a chosen User-Agent, but that changes only one visible surface. Sites may also compare fonts, GPU and CPU characteristics, IP address, TLS details, browser APIs, timing, scrolling, typing, and other behavior. Research findings about particular agents and detectors are controlled-study results; they are not guarantees about every browser, website, or future detector.
1. What is a browser agent?
A web user agent is software that interacts with websites on a user’s behalf. That includes a traditional browser, an accessibility tool, an automation script, or an AI system that navigates pages and performs requested actions. The W3C definition covers software that renders content as well as software that performs actions authorized by a user (W3C Web User Agents).
An agent’s identity has at least three layers:
| Layer | What it means | Examples |
|---|---|---|
| Declared identity | Values sent openly by the client | User-Agent, Accept headers, Client Hints, cookies |
| Technical fingerprint | Properties inferred from the browser, device, and connection | Fonts, GPU, CPU, rendering differences, IP, TLS |
| Behavioral profile | How requests and interactions occur | Typing cadence, scrolling, mouse movement, navigation timing |
These layers can be combined for security, fraud prevention, abuse controls, compatibility decisions, or tracking. A fingerprint can help authenticate a session, but it is not proof that an action is authorized.
2. What is browser fingerprinting?
Browser fingerprinting is the collection and combination of observable properties to distinguish one client from another. WebKit lists fonts, User-Agent strings, GPU and CPU details, IP address, and TLS connections among relevant vectors (WebKit Tracking Prevention Policy).
A site might collect a User-Agent header on the server, then use JavaScript to inspect supported features or rendering behavior. It can compare those results with network information and the sequence of actions that follows. The resulting record may be used only for a short-lived security decision or retained to link visits over time.
Fingerprinting has legitimate security uses, but it has privacy costs. The W3C warns that a fingerprint can enable tracking without clear controls and is usually difficult for a user to clear or reset (W3C fingerprinting guidance).
Fingerprinting is probabilistic
Signals are noisy. Many people can share the same browser version, operating system, screen size, or cloud IP range. A site therefore combines signals and assigns a risk or similarity judgment rather than reading a universal “agent ID.” The same agent can also look different after a browser update, network change, privacy setting, or device migration.
3. How do websites know what browser I’m using?
The first clue is usually the HTTP request. A browser sends headers such as User-Agent, Accept, and language preferences. A server can log these values before returning a page.
After the page loads, client-side code may inspect capabilities and platform properties. Rendering tests can reveal implementation differences in graphics, fonts, canvas, audio, or layout. Sites can also observe connection information, cookies, cache state, and request timing. None of these signals alone identifies a person reliably; their combination can make a session more distinguishable.
What is a User-Agent string?
A User-Agent string is an HTTP request header that describes the client software and often its operating-system family. It is historical compatibility metadata, not a cryptographic identity. User-Agent sniffing is brittle because a browser name does not guarantee that a feature exists. MDN recommends feature detection where possible (MDN browser detection guidance).
curl -I https://example.com
# Look for a response such as:
# user-agent handling is based on the User-Agent request header
To inspect the request you send:
curl -v https://example.com -o /dev/null
Changing this header may alter compatibility behavior, but it does not make the rest of a session resemble another browser.
4. User-Agent Client Hints
User-Agent Client Hints move some identity data from passive disclosure toward explicit requests. An origin can send Accept-CH to request selected hints under the framework defined by RFC 8942. Chrome’s guidance recommends minimizing User-Agent data and using feature detection, progressive enhancement, or responsive design when those solve the problem (Chrome User-Agent Client Hints).
HTTP/1.1 200 OK
Accept-CH: Sec-CH-UA, Sec-CH-UA-Platform
HTTP/1.1 200 OK
Sec-CH-UA: "Chromium";v="140"
Sec-CH-UA-Platform: "Linux"
Client Hints do not eliminate fingerprinting. High-entropy values can increase linkability, and requested hints still disclose information. Ask only for data needed for a documented compatibility or security function.
5. Can websites detect AI browser agents?
They can sometimes classify an agent as automated or unusual, but detection is not universal and no single test is conclusive. Detection systems compare several layers:
- Network: IP reputation, hosting ranges, connection patterns, and TLS characteristics.
- HTTP: headers, cookies, request order, missing browser fields, and Client Hints.
- Browser: automation-visible properties, rendering behavior, and API consistency.
- Behavior: typing, scrolling, pointer movement, dwell time, navigation paths, and repeated actions.
A 2026 preprint studying six LLM-based web agents reported that combined network, HTTP, and browser signals distinguished the tested agents from human traffic and from one another. A separate controlled study of seven AI browsing agents found more distinctive typing, scrolling, and mouse behavior across its tested tasks. Those results describe the agents, sites, and defenses in each experiment; they are not market-wide detector accuracy.
Another 2026 benchmark reported that two binary classifiers misclassified 39.1% and 34.5% of tested AI agents as human. The figures apply only to that benchmark and classifier setup. Adding an explicit agent class changed the outcome, which illustrates why labels and evaluation design matter.
See the reported studies: multi-layer fingerprinting, FP-Agent, and behavioral detection.
6. A safe way to inspect your own agent
Use inspection to document compatibility and privacy exposure on systems you own or are authorized to test. Do not treat these checks as a way to defeat a particular site’s controls.
Browser-side JavaScript
const identity = {
userAgent: navigator.userAgent,
platform: navigator.platform,
language: navigator.language,
languages: navigator.languages,
hardwareConcurrency: navigator.hardwareConcurrency,
deviceMemory: navigator.deviceMemory,
maxTouchPoints: navigator.maxTouchPoints,
screen: {
width: screen.width,
height: screen.height,
pixelRatio: window.devicePixelRatio
},
timezone: Intl.DateTimeFormat().resolvedOptions().timeZone,
webdriver: navigator.webdriver
};
console.log(identity);
Some properties may be unavailable, reduced, or deliberately generalized by the browser. Treat the output as a compatibility snapshot, not a permanent identifier.
Playwright with an explicit User-Agent
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext({
userAgent: 'Mozilla/5.0 (compatible; DocumentationBot/1.0)',
viewport: { width: 1440, height: 900 }
});
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.locator('body').innerText());
await browser.close();
Playwright documents User-Agent configuration as a browser-context option (Playwright Browser API). It changes a configurable surface only; it does not change every network, rendering, or behavioral signal.
7. Can I stop my browser from being fingerprinted?
You can reduce unnecessary exposure, but no general setting makes every browser indistinguishable. Practical steps include:
- Keep your browser current and use its built-in anti-tracking protections.
- Limit extensions; unusual extension combinations can add identifying variation.
- Block unnecessary third-party scripts and storage where that does not break a site you trust.
- Prefer browsers or modes that deliberately standardize common values.
- Separate work profiles from personal profiles when you need different cookies and sessions.
- Review permissions for location, camera, microphone, notifications, and persistent storage.
- Use feature detection in your own applications instead of collecting browser names for compatibility.
There are tradeoffs. Aggressive blocking can break authentication, payments, accessibility features, or fraud controls. A highly unusual configuration can itself become distinctive. W3C guidance recommends exposing only the entropy required for a function and making access visible or opt-in where possible.
8. Designing detection responsibly
If you operate a service, document what you collect and why. Separate compatibility logic from security classification, and measure false positives for legitimate browsers, accessibility tools, privacy browsers, and authorized agents.
| Decision | Recommended practice |
|---|---|
| Compatibility | Use feature detection, progressive enhancement, and responsive design. |
| Abuse prevention | Combine multiple signals and provide a recovery path for false positives. |
| Privacy | Collect the minimum entropy, set retention limits, and explain the purpose. |
| Authorization | Require user confirmation for consequential actions; a fingerprint is not authorization. |
| Evaluation | Test across browsers, devices, networks, and agent types; report the study scope. |
Authenticated agents require extra care. Chrome’s WebMCP security guidance describes malicious instructions hidden in tool definitions and malicious content returned by otherwise trusted sites. Keep a human in the loop for state-changing operations and treat tool results as untrusted input (Chrome agent security guidance).
9. Capture a reproducible view of an agent session
Screenshots help you compare what an agent saw after a navigation, consent flow, or feature-detection step. A local browser gives maximum control:
import { chromium } from 'playwright';
const browser = await chromium.launch();
const context = await browser.newContext({
viewport: { width: 1440, height: 900 },
colorScheme: 'light',
locale: 'en-US'
});
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle' });
await page.screenshot({ path: 'agent-view.png', fullPage: true });
await browser.close();
For repeatable evidence, record the URL, timestamp, viewport, User-Agent policy, locale, timezone, and whether the page reached a stable state. Avoid saving secrets or personal data in screenshots and logs.
10. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', data);
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, User-Agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work to simplify migration.
An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so an AI agent can request captures through its tool connection.
Plans include 1,000 shots per month free with no card; paid plans start at $5 for 3,000 shots. Yearly billing provides two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The site shows the wrong browser | User-Agent sniffing or a stale cache | Use capability checks, inspect Client Hints, and invalidate the relevant cache. |
| An agent is challenged or blocked | The service classified combined network, browser, or behavior signals as risky | Use the site’s documented access path, reduce request rate, and provide a recovery or human-review path. |
| Changing User-Agent did not help | Other signals still differ | Treat User-Agent as compatibility metadata, not an identity reset. |
| Client Hints are missing | The origin did not opt in, the browser limited them, or the request is not eligible | Send Accept-CH only when needed and keep a feature-detection fallback. |
| Screenshot is blank | The page timed out, failed to load, or requires a later wait condition | Wait for a selector or network idle, verify the URL, and inspect the page verdict. |
| Consent UI appears in a capture | The banner uses an unsupported flow or cleanup was disabled | Handle the consent step in your browser workflow or configure the relevant ScreenshotNeo option. |
| Different runs produce different fingerprints | Browser updates, fonts, IPs, timing, or session state changed | Pin the environment, record configuration, and compare only like-for-like runs. |
12. Performance, reliability, and cost
- Performance: Fingerprint collection adds client-side work and network requests. Request only signals needed for a concrete decision. For screenshots, reuse browser contexts or use an API when starting a full browser for every URL is expensive.
- Reliability: Network idle is not always a meaningful readiness signal for applications with continuous connections. Prefer a stable selector or application-specific readiness marker, and capture verdict metadata.
- Privacy: High-entropy data improves distinction but increases linkability. Set retention and access controls for logs.
- Cost: Local automation consumes compute and bandwidth. ScreenshotNeo bills only clean shots; bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Cache TTLs, bulk capture, and asynchronous jobs can reduce repeated work.
13. FAQ
Is fingerprinting the same as cookies?
No. Cookies are stored identifiers. Fingerprinting derives identity from observable properties. Sites may use both.
Does private browsing make fingerprinting impossible?
No. It can change storage and session behavior, but websites may still observe browser, device, network, and interaction signals.
Can a fingerprint prove that traffic came from an AI?
No. It can support a classification decision. A result depends on the signals, model, threshold, and evaluation environment.
Should I block every automated browser?
Usually not. Accessibility tools, monitoring jobs, testing systems, and authorized agents can be legitimate. Use the least disruptive control that addresses the abuse you can demonstrate.
What should an AI agent disclose?
Follow the site’s policy and the user’s authorization. An honest, stable identity and clear confirmation for consequential actions are safer than pretending that one configurable header represents the whole client.
Start with 1,000 free ScreenshotNeo screenshots each month, with no card required.


