ScreenshotNeo

BlogGuides

How Bot Detection Works and How to Test Your Website Against Bots

Learn how layered bot detection works, then safely test your own site with a controlled workflow that protects legitimate users and automation.

By the ScreenshotNeo team4 October 20269 min read

Bot detection estimates whether a request is automated by combining signals such as network reputation, request patterns, session behavior, endpoint context, and browser checks. The site then decides what to do: allow, log, challenge, rate-limit, delay, or block. To test your own site, start with a route-specific threat model, establish a baseline, send a small amount of labeled traffic in a controlled environment, and verify both abusive patterns and legitimate clients before enforcing rules.

A bot is not automatically malicious. Search crawlers, uptime monitors, accessibility tools, mobile apps, and API clients may all be expected automation. The goal is to reduce harmful activity while preserving traffic your site needs. OWASP recommends matching controls to abuse risks and applying them across edge, application, and business-logic layers. OWASP Bot Management and Anti-Automation Cheat Sheet

How bot detection works

Detection systems combine observations into an estimate. No single signal reliably identifies every bot, and vendors expose different signals and scoring methods.

Signal What it can indicate Limit to account for
IP and network reputation Requests associated with known proxies, hosting networks, or prior abuse. Shared networks and VPNs can include ordinary users; IP-only policies can also be evaded by distributed traffic.
Headers and request fingerprints Missing, inconsistent, or repeated request characteristics. Legitimate API clients and older software may send unusual headers.
Rate and velocity Repeated requests, login failures, or unusually fast actions over a window. Busy users, integrations, and bursty workloads can look unusual; use route and identity context.
Session and identity behavior Patterns such as many account attempts from related sessions or repeated account creation. Shared devices and privacy tools complicate identity assumptions.
Endpoint and business context Whether behavior is risky for a particular operation, such as login or checkout. A rate acceptable for a catalog may be unsafe for a credential or payment endpoint.
Browser-side checks Whether a browser executes a script or completes a challenge. Disabled JavaScript, network errors, and assistive or privacy software can affect results.

Some products combine signals into a score. Cloudflare Enterprise Bot Management, for example, assigns a per-request score from 1 to 99; its documentation describes score bands for definite and likely automation. These are Cloudflare product values, not a universal bot score. Cloudflare Bot Management scoring

Detection and response are separate decisions. A low-confidence signal can be logged for analysis; a stronger pattern may justify a challenge or rate limit; a high-confidence, well-understood abuse pattern may justify blocking. OWASP recommends layered controls and proportional responses rather than treating one signal as proof.

Start with a route-specific threat model

List the routes that matter, the abuse you expect, and the legitimate clients that must continue to work. OWASP identifies examples including credential stuffing, scraping, inventory hoarding, fake accounts, card testing, fake reviews, and click fraud. Assign controls to the relevant endpoint instead of applying one blanket policy to every request.

Route or function Possible abuse Legitimate traffic to preserve
Login Repeated password guesses or credential stuffing Normal users, password managers, support workflows
Signup Fake account creation New users, approved identity providers
Catalog or search Excessive scraping or inventory hoarding Search crawlers, accessibility tools, product integrations
Checkout Card testing, scalping, or abusive purchase attempts Customers and payment-provider callbacks
Public API Excessive queries, scraping, or resource exhaustion Partners, mobile apps, documented API clients

For each route, write down the signal you expect to observe, the action you would take, and how you will recognize a false positive. Prefer limits keyed to meaningful dimensions—such as endpoint, session, or authenticated identity—over relying only on source IP.

Test bot detection safely

  1. Set authorization and scope. Use staging where possible. For production, get approval for the exact routes, test window, request volume, and rollback owner. Do not probe third-party sites, real accounts, or data you do not control.
  2. Record a normal baseline. Note traffic volume, status codes, route distribution, login failures, challenge rates, and expected crawler, monitoring, API, and mobile traffic. If your provider has analytics or security events, capture the current view before changing rules.
  3. Label test traffic. Use a dedicated test account and an identifying user agent or other agreed marker. Begin with a few requests at a low rate. Never use real credentials or attempt to evade a control.
  4. Change one behavior at a time. For example, compare an ordinary request with a request missing a header, a modest rate increase, or repeated failed logins against the test account. These are controlled test examples, not recommended universal thresholds.
  5. Observe before enforcing. Use a log, preview, or monitor action if available. Inspect matched rules, event records, response status, challenge presentation, errors, and latency.
  6. Test known-good clients. Exercise normal browser journeys and relevant crawlers, uptime monitors, partner APIs, mobile clients, and accessibility software. Verify crawler identity using the approach appropriate to your provider rather than trusting a user-agent string alone.
  7. Tune and repeat. Change one threshold or rule at a time. Compare intended detections, false positives, challenge completion, latency, and business outcomes.
  8. Keep rollback ready. Save the previous configuration and decide who can disable the rule if legitimate requests fail. Recheck after application, traffic, or provider changes.

Cloudflare recommends reviewing bot analytics and Security Events, targeting rules to observed patterns, and tuning based on results. Detailed analytics and available controls depend on the product and plan. Cloudflare guidance for stopping malicious bots while allowing legitimate traffic

Choose a response that matches confidence

Response When it fits What to monitor
Log or monitor Early rollout, weak signal, or a rule whose impact is not yet known. Match volume and whether known-good clients are included.
Allow or create an exception Verified crawlers, internal APIs, partners, and monitoring tools. Keep exceptions narrow and review whether they still apply.
Challenge Suspicious requests where an interaction can distinguish some automation from people. Completion rates, accessibility impact, and abandoned journeys.
Rate-limit or delay Excessive behavior where slowing requests reduces harm while preserving some access. Legitimate bursts, retries, and downstream load.
Block High-confidence abuse with a clear route-specific rule and tested exceptions. False positives, support reports, and changes in attack behavior.

Do not assume a challenge works for every client. Cloudflare notes that its simple Bot Fight Mode applies across a domain and may challenge API or mobile traffic. Verify the scope and behavior of any control before enabling it. Cloudflare Bot Fight Mode

Browser checks are evidence, not proof

Client-side JavaScript detection can add useful context, but it has important boundaries. Cloudflare documents that its JavaScript Detection script is injected into HTML responses, does not cover AJAX calls, and populates a cookie field for later rules. A separate rule must act on a failed result. The documentation advises against applying that rule to the first request or traffic that does not expect browser JavaScript.

A failed result may have benign causes, including disabled JavaScript or network problems. Treat it as one input alongside endpoint, session, and request context. Do not use a browser-only check as the sole gate for APIs or other clients that do not run page scripts. Cloudflare JavaScript Detections documentation

What to compare in bot protection

  • Visibility: Can you inspect events, scores, reasons, and useful analytics?
  • Scope: Are controls domain-wide or configurable by route and endpoint?
  • Actions: Can you log, allow, challenge, rate-limit, delay, or block?
  • Known-good clients: Can you handle verified crawlers and make narrow exceptions for APIs, monitoring, partners, and mobile apps?
  • Operations: How much rule tuning and false-positive investigation will the team need to do?
  • Privacy and access: What client signals are collected and retained? Can challenges create barriers for people using assistive technology?
  • Deployment and plan: Does protection require a particular edge provider, and can it affect cached or static content?

Cloudflare documents a tradeoff between domain-wide Bot Fight Mode and more granular Enterprise Bot Management controls. Check the provider’s current plan details and behavior before purchasing or deploying a feature. Cloudflare bot documentation

Troubleshooting common test results

Symptom Likely cause What to do
Expected test request is not detected The signal is not covered by the rule, the rule is scoped to another route, or the provider has not observed the context it needs. Check event logs and rule scope. Confirm the request reached the protected layer, then vary one test attribute at a time.
Ordinary browser or API traffic is challenged A broad rule, shared signal, or domain-wide action is catching legitimate clients. Review matched events, narrow the route or condition, and add a verified, narrowly scoped exception before enforcement.
JavaScript detection appears to fail on the first visit The HTML request needed to run the script may not have occurred yet, or the rule is being applied too early. Follow the provider’s documented sequencing; do not use the result as a first-request verdict.
AJAX or API calls lack a browser signal The documented JavaScript check may only be populated through HTML responses and may exclude AJAX. Use API-appropriate controls such as identity, route context, and rate limits rather than requiring a browser cookie.
Monitor or partner integration stops working A legitimate client was not identified or an enforcement action has wider scope than expected. Confirm its identity through an appropriate method, add a narrow exception, and repeat the normal journey.
Users report inaccessible or confusing challenges The challenge may not work with a client, assistive technology, or the user’s network conditions. Review challenge outcomes and provide an accessible route through the flow; consider a less disruptive response for that route.
Rule changes cause an outage or unexpected failures The rule was enforced without a tested rollback or included a critical route. Use the saved prior configuration or disable the affected rule, then inspect events and retest in monitor mode.

Performance, reliability, and cost

Bot controls can add request processing, browser challenges, or extra round trips. Measure latency and completion on the actual routes and clients you protect; do not assume a challenge or script has negligible impact. Rate limits also need to account for legitimate bursts and retries.

Reliability depends on correct scope, exceptions, observability, and a rollback path. Revisit rules when routes, authentication, client behavior, or traffic patterns change. Keep application-level protections for business-critical operations even when an edge provider filters traffic, since no single detection layer sees every signal.

Costs and feature availability vary by vendor, product, and plan. Compare the controls and analytics you will actually use, operational effort, and impact of false positives. Confirm current prices and plan limits with the provider; this guide does not assume a particular plan includes a given feature.

Or skip the browser setup

If you need a clean screenshot of a page while documenting bot behavior or reviewing a test route, ScreenshotNeo is a website screenshot API and MCP server. It captures pages as PNG, JPEG, WebP, or PDF. A screenshot is useful for observing a page’s visible result; it does not determine whether a request is a bot or replace your security logs.

One GET request returns the capture. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, and failed loads are never billed. Response headers report the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.

Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.

FAQ

Does a high request rate prove that traffic is a bot?

No. It is a signal to evaluate in context. Legitimate clients can generate bursts, and harmful automation can spread requests across addresses or sessions.

Should every site block bots?

No. Many sites depend on crawlers, monitoring, APIs, and other automation. Decide which behavior is harmful on each route and preserve expected clients.

Can JavaScript detection identify every automated client?

No. It only supplies a limited browser-side signal and can fail for benign reasons. Combine it with other evidence and use controls suitable for non-browser clients.

What is the safest first enforcement action?

There is no universal action. Start by observing the rule, verify its matches and exceptions, then choose the least disruptive response that addresses the route’s risk.