ScreenshotNeo

BlogGuides

How Bot Detection Works and How to Test Your Website

Learn how bot detection combines request and browser signals, how to verify claimed crawlers, and how to test controls without blocking legitimate traffic.

By the ScreenshotNeo team4 October 202610 min read

Bot detection estimates whether a request is automated by combining evidence such as request patterns, headers, session behavior, and browser signals. It does not prove intent: legitimate crawlers, monitoring services, accessibility tools, and user-directed agents are also automated. To test your own site, map likely abuse by endpoint, exercise authorized human and automated flows, observe logs before blocking, and tune controls against both abuse and false positives.

There is no universally reliable signal or score threshold. A User-Agent can be copied, browser checks can be bypassed, and legitimate clients may look unlike a typical browser. Treat detection as an input to a proportionate decision—allow, log, rate-limit, challenge, or block—not as a verdict on its own. [Cloudflare’s bot detection engines] [OWASP Bot Management and Anti-Automation Cheat Sheet]

How bot detection works

A detector looks for evidence that a request was generated or controlled by automation. Simple systems match known patterns. More involved systems combine signals across requests, sessions, and browser execution, then estimate how likely a request is to be automated. The available evidence and the cost of a mistake vary by route.

Cloudflare documents one vendor implementation with a heuristic engine that checks requests against patterns and fingerprints, optional JavaScript detections that can identify headless browsers and other fingerprints, and a machine-learning engine that considers request features such as headers, session characteristics, and browser signals. This is an example, not a description of every bot-detection system. [Cloudflare: bot detection engines]

Signals are evidence, not proof

  • Request details: headers, User-Agent, request method, path, source network, and consistency between claimed client details. These are easy to observe, but some are easy to forge.
  • Request patterns: repeated requests, unusual velocity, systematic enumeration, or behavior concentrated on sensitive routes.
  • Session and browser behavior: whether a client maintains expected session state or executes browser-side checks. Such checks add evidence but can affect APIs, mobile apps, privacy tools, and assistive technology.
  • Known identities and validation: verified crawler records, published IP ranges, reverse DNS, or signed bot identity mechanisms where supported.

Signals can conflict. A real user behind a shared network may resemble a high-volume client; a scraper can spread requests across addresses. Scores summarize evidence and should be interpreted in context.

For example, Cloudflare Bot Management documents scores from 1 to 99 and says scores below 30 are commonly associated with bot traffic. That is Cloudflare product guidance, not an industry standard or a universal blocking threshold. [Cloudflare: bot score]

Decide what to protect before testing

Start with the endpoint and the harm you want to prevent. A login form, public search page, checkout, and API have different risks and legitimate clients. OWASP recommends layered defenses across the edge, application, and business logic, with controls matched to the use case. [OWASP guidance]

Endpoint or flow Possible abuse Useful controls to evaluate
Login Credential stuffing or repeated password guessing Rate limits by account and source, anomaly monitoring, step-up verification, account protections
Signup Fake accounts, automated submissions, verification abuse Velocity limits, risk-based verification, signup quality monitoring
Search or catalog Scraping, enumeration, costly query patterns Per-identity or per-session limits, query controls, selective challenge
Checkout or scarce inventory Card testing, scalping, inventory hoarding Purchase limits, payment risk controls, queues, velocity rules
Public API Abusive usage, probing, credential misuse Authentication, per-key or per-identity quotas, route-specific limits, structured logs

These are starting points, not prescriptions. Consider the cost of blocking a real customer or partner alongside the cost of abuse. A broad domain-wide challenge can disrupt mobile clients and APIs that do not behave like browser sessions.

How to test bot detection on your website

  1. Inventory important routes. Record the endpoint, expected human flows, legitimate automated clients, likely abuse, and the business impact of a false positive. Include APIs, mobile applications, accessibility flows, monitoring, and search crawlers where relevant.
  2. Define a controlled test matrix. Exercise normal human use, known legitimate automation, and authorized simulations of the abuse patterns relevant to each route. Use a staging environment when practical. Keep test volume bounded and within your authorization; this workflow is not a claim that a penetration test or benchmark has been conducted.
  3. Establish a baseline in observe or log mode. Capture route, time, action, available score or signals, identity dimension used for limits, and outcome. Check whether events correspond to known legitimate services before changing enforcement.
  4. Test each response separately. Where your platform supports them, compare allow, log, rate-limit, challenge, and block on the specific routes that need protection. Confirm that the action occurs at the intended layer and does not unintentionally cover APIs or static resources.
  5. Measure both sides of the error. Track abuse that still succeeds, false positives, challenge completion or abandonment where available, latency, support reports, and impact on legitimate clients. A detector that catches more automation may still be a poor choice if it blocks important users.
  6. Tune and repeat. Adjust one rule or threshold at a time, keep decision logs, and repeat the same human and automation cases after each change. Avoid hard-blocking on a weak single signal.

Cloudflare documents using analytics and logs to analyze patterns and tune rules. Its domain-wide Bot Fight Mode can challenge API or mobile-app traffic; its troubleshooting material also notes that testing and monitoring tools with bot-like User-Agent strings may be flagged. Test those cases explicitly before enabling broad actions. [Cloudflare bot solutions] [Cloudflare bot troubleshooting]

A compact test matrix

Case What to verify
Ordinary browser user Expected route works without an unnecessary challenge or delay
Authorized monitor or API client Known automation receives the intended response and is logged accurately
Legitimate crawler Identity is validated using an authoritative method; crawl access matches policy
High-rate authorized simulation Limits trigger at the intended identity and route scope
Invalid or inconsistent client Signals are logged and the selected response is proportionate
Accessibility and mobile flows Users and apps have a usable path when a challenge or script check is applied

Do not use someone else’s site as a target for these tests. Keep simulations within systems and accounts you control or have explicit authorization to assess.

How to verify a request claiming to be Googlebot

A User-Agent string is self-asserted and can be copied. For requests claiming to be Google, Google recommends verifying the source with reverse DNS or checking its IP against the published crawler and fetcher IP ranges. Identify which Google request category is involved—common crawler, special-case crawler, or user-triggered fetcher—because their policies can differ. [Google: verify Googlebot and other Google crawlers]

  1. Record the source IP and claimed User-Agent from your server or edge logs.
  2. Perform the verification method described in Google’s current documentation: reverse DNS validation or a match against the published IP ranges.
  3. Use the verified identity and request category in your allow policy; do not allow traffic solely because the User-Agent contains “Googlebot.”
  4. Recheck the authoritative records and your verification implementation as crawler guidance or address ranges change.

Google describes Web Bot Auth support as experimental, based on an IETF draft, and says not all of its user agents use it or sign every request. Google advises site owners to continue relying on IP addresses, reverse DNS, and User-Agent strings during rollout. Treat signed bot identity as an emerging signal, not a universal replacement. [Google crawler verification guidance]

Cloudflare’s verified-bot model similarly describes honest, deterministic self-identification and non-abusive behavior as criteria, with methods including Web Bot Auth and IP validation. Verification establishes identity more than intent: a verified service can still make requests that your endpoint policy should constrain. [Cloudflare: verified bots]

Choose controls and scope them carefully

When comparing bot-management approaches, evaluate the detection signals you can inspect, where the control applies, which actions are available, how rules can be tuned, false-positive effects, privacy and retention, and operational effort. Cloudflare documents options ranging from baseline Bot Fight Mode to more granular Super Bot Fight Mode and Enterprise Bot Management. OWASP’s endpoint-specific guidance is useful regardless of vendor. [Cloudflare bot solutions] [Cloudflare bot management] [OWASP cheat sheet]

  • Scope rules to the route and method where the risk exists instead of challenging the entire domain by default.
  • Apply limits across appropriate dimensions: source IP, account or identity, API key, endpoint, and time window. Shared IP limits alone can penalize users behind carrier or corporate networks.
  • Use business rules for business risks: for example, purchase limits for scarce items or per-account controls for repeated actions.
  • Offer accessible alternatives to user-facing challenges and preserve legitimate API and mobile paths.
  • Log the signals and action that drove each decision, while retaining only the data needed for security and operations.

Performance, reliability, and cost considerations

  • Latency: additional browser-side checks, challenges, or external lookups can add work to a request. Measure route latency and completion rates in your own environment; this dossier provides no benchmark.
  • Availability: a detector or challenge provider can become a dependency in the request path. Decide what should happen if a check is unavailable, especially for login, checkout, and APIs, and monitor the fallback behavior.
  • False positives: tune against known crawlers, monitoring, API clients, mobile apps, and assistive technology. Keep a way to identify and resolve legitimate blocked traffic.
  • Operational cost: account for service plan eligibility, engineering and support time, challenge friction, log storage, and the business cost of abuse that remains. Product-specific pricing is not covered by the cited bot-management documentation.
  • Privacy: document which request and browser signals are collected, who can access them, and how long they are retained. Keep observability sufficient to explain decisions without retaining unnecessary data.

Or skip the browser setup

If you need a page screenshot while investigating a suspected bot check or rendering issue, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. It can accept cookie and consent banners as a visitor and remove 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which outcome occurred. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents.

Install the Python dependency with python -m pip install requests, set your API key, then run:

import os
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": os.environ["SCREENSHOTNEO_API_KEY"], "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)

For equivalent requests, see the ScreenshotNeo API documentation. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const data = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. The MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.

Troubleshooting common problems

Symptom Likely cause What to check or change
Googlebot-looking traffic is blocked The rule trusts a User-Agent or rejects a real crawler based on a weak signal Verify the source with Google’s reverse-DNS or published IP-range guidance, identify the crawler category, then update a narrowly scoped rule
API or mobile requests receive challenges A broad browser-oriented bot control is applied domain-wide Inspect the affected route and rule action; exclude or separately protect known API and app paths where appropriate
Monitoring checks are flagged The monitor uses a bot-like User-Agent, request rate, or browser profile Confirm the monitor’s source and identity, then create a scoped exception or adjust its request cadence; retain monitoring visibility
Real users behind one network are rate-limited Limits use shared source IP as the only identity dimension Combine IP with account, API key, session, or endpoint context and reassess the window and threshold
Abusive traffic passes despite a high-confidence rule The rule covers the wrong route or identity dimension, or the actor changes behavior Review logs by endpoint and account, add business-logic controls, and test the changed pattern rather than only repeating the original case
Challenges reduce signups or purchases Legitimate users encounter friction or cannot complete an inaccessible challenge Review user impact and completion data, narrow challenge scope, and provide an accessible alternative
Logs do not explain a block Decision signals or rule actions are not recorded at a useful level Enable decision logging, include route and action, and define privacy-aware retention

Frequently asked questions

Does a browser automation tool prove a visitor is malicious?

No. Automation describes how a request is produced, not whether it is abusive. Judge it against endpoint policy, identity, behavior, and authorization.

Should every bot be blocked?

No. Search crawlers, monitoring, accessibility tools, and user-directed agents can be useful. OWASP frames the goal as making abusive automation more costly while preserving legitimate users and bots.

Can I rely on a bot score alone?

No. A score is a system-specific estimate. Combine it with route context, verified identity where available, business rules, and observed impact.

Is Web Bot Auth ready to verify every crawler?

No. Google describes its support as experimental and says not all Google user agents sign every request. Continue using the documented IP and reverse-DNS methods during rollout.

Sources