How Websites Detect and Block Web Scraping
Websites combine request, behavior, and browser signals to identify automated traffic. Learn how detection works, what robots.txt can do, and how to choose proportionate defenses.
Websites detect and block scraping by combining signals from incoming requests, browser behavior, traffic patterns, and sometimes client-side JavaScript. Detection informs a policy: allow the request, block it, ask for a challenge, or limit how often an operation can be repeated. There is no single signal or threshold that identifies every scraper, and automated traffic is not always harmful.
If you operate a site, start by deciding which routes and automated clients should be allowed, then apply proportionate controls and monitor their effect on real visitors and APIs. If you are building a compliant data workflow, use an authorized API or obtain permission; do not treat detection signals as a checklist for bypassing a site’s controls.
1. What bot detection systems look at
Bot detection is layered. A provider can combine known fingerprints and heuristics with machine learning, behavioral patterns, traffic baselines, and client-side JavaScript signals. The exact mix depends on the provider, product, configuration, and plan. Cloudflare says it uses multiple detection engines because different bot types need different strategies; its documentation is an example of one vendor’s capabilities, not a universal blueprint.
| Signal family | What it can indicate | Important limitation |
|---|---|---|
| Known signatures and heuristics | A request or client resembles a known automated pattern. | Signatures cover known patterns; a match is not proof of intent. |
| Request behavior | Request frequency, repetition, or route sequences differ from expected use. | Busy legitimate users and integrations can also generate unusual patterns. |
| JavaScript signals | Client-side checks can add information about browser execution. | These checks may not suit every API or client and should be scoped carefully. |
| Traffic baselines and broader patterns | Activity can be assessed against patterns observed across a site or network. | Baselines vary, and shared networks can make attribution uncertain. |
| Machine learning and behavior analysis | Multiple signals can be evaluated together to classify traffic. | Classification is probabilistic and can produce false positives or negatives. |
For example, Cloudflare documents scraping-specific analysis using zone-level traffic patterns and dynamic analysis by ASN and JA4 fingerprint. It says matches are recalculated, rather than treating one fingerprint as a permanent flag. These are Cloudflare-specific examples; other providers may use different signals or expose different controls.
Cloudflare also documents a bot score from 1 to 99. In its system, scores below 30 are commonly associated with bot traffic. That is a Cloudflare scale and threshold, not an industry standard, and a score does not prove that a request is scraping.
2. What a site can do after detection
Detection is separate from enforcement. A site or its protection provider uses rules to decide what action to take, often based on a route, request class, or other configured condition.
| Action | Purpose | Tradeoff |
|---|---|---|
| Allow | Permit expected traffic, including useful or verified crawlers. | Allowed traffic can still consume resources; monitor sensitive operations. |
| Block | Deny traffic that meets a sufficiently strong rule. | A broad rule can reject legitimate users or integrations. |
| Challenge | Ask a suspicious visitor to complete an additional check. | Challenges can interrupt real visitors and break API clients that cannot complete them. |
| Rate-limit | Cap repeated requests or operations within a defined period. | Limits that are too strict can interfere with normal bursts or shared users. |
Cloudflare documents challenge pages and JavaScript detections as tools that can be used in security rules. Its rate-limit guidance gives repeated price lookups as an example of an operation an operator might limit to make large-scale catalog extraction harder. These controls are most useful when aimed at sensitive routes and operations, with monitoring for effects on genuine use.
Automated traffic is not automatically unwanted. Search crawlers and other bots may help a site, while other automation may harm it. Define which clients and behavior are beneficial, then create rules that reflect those distinctions instead of treating every automated request alike.
3. What robots.txt can and cannot do
robots.txt communicates crawler preferences. Google says Googlebot and other respectable crawlers follow those instructions, while other crawlers might not. It is useful for telling compliant crawlers which paths to avoid; it is not access control and cannot stop a client that chooses to ignore it.
If a resource must be protected, enforce access on the server with authentication or authorization, application rules, a web application firewall (WAF), rate limits, or another control suitable for that resource. Do not put secrets in a path and assume that disallowing the path in robots.txt makes it private.
4. Choosing controls for your site
- Inventory routes and clients. Identify public pages, APIs, login and account paths, expensive operations, and known useful crawlers or integrations.
- Choose the signal and action for each case. For a high-cost operation, a route-specific rate limit may be a better first control than a site-wide challenge. For clearly disallowed access, use an authorization check or a narrowly scoped block.
- Keep APIs usable. If a challenge is unsuitable for API clients, exclude the relevant API paths from challenge rules and protect them with controls those clients can satisfy.
- Roll out and observe. Review challenge rates, blocked requests, rate-limit events, support reports, and API failures. Adjust rules when legitimate traffic is affected.
- Revisit policy and provider settings. Detection engines, rule availability, and controls vary by provider and plan. Confirm the features available in your own configuration.
When comparing managed options, use the same questions for each provider: what signals are available, which actions can rules take, how precisely can they target routes or clients, what monitoring and tuning do they require, and which product tier includes them? Cloudflare and Google Cloud Armor document managed bot controls, but the available sources do not establish an independent cross-vendor efficacy ranking.
5. A practical example: rate-limit an expensive operation
Consider a product catalog where repeated price lookups are costly. The defensive approach is to identify the actual lookup route, set a limit that fits normal usage, and monitor whether legitimate visitors or integrations are affected. The exact rule syntax and suitable threshold depend on your application and provider; no universal value is appropriate for every site.
Policy outline (provider-neutral)
route: /api/products/{product_id}/price
measure: repeated price-lookup requests over a defined time window
action: apply a limit appropriate to expected legitimate use
exceptions: explicitly authorized clients, if needed
monitor: rate-limit events, successful lookups, user reports, API errors
adjust: tune the scope and limit when normal use is affected
This outline is not copy-and-paste configuration for a particular WAF. Use your provider’s current documentation to implement the rule, and keep authentication and authorization checks in the application where access depends on identity or entitlements.
6. If your authorized task is taking website screenshots
A screenshot workflow should use pages you are authorized to capture and should respect the site’s access controls. When a request encounters a challenge, bot check, blank page, timeout, or failed load, treat that as a failed capture or a signal to resolve access with the site owner—not as an invitation to evade the control.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. Its clean-shot flow accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status.
For an authorized page, this cURL request saves a WebP screenshot. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python equivalent:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js equivalent:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo also provides an MCP server for AI agents, including Claude, Cursor, and other MCP clients. Its tools include take_screenshot, get_page_info, and capture_pdf. The API supports full-page and element captures, viewport and device settings, custom CSS and JavaScript, waits, request blocking, cookies and headers, caching, signed links, asynchronous jobs, bulk capture, and more. Plans include 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
7. Troubleshooting detection and mitigation
| Symptom | Possible cause | What to check |
|---|---|---|
| Legitimate visitors are challenged | A broad rule or a classification signal is catching normal traffic. | Review the affected route and rule conditions; narrow scope and monitor after changes. |
| An API client receives a challenge page | A browser-oriented challenge rule includes an API path. | Exclude paths that should not receive interactive challenges and use API-appropriate controls. |
| Rate limits affect normal bursts | The scope or limit does not match real usage, including shared clients. | Review the operation and traffic distribution; tune the rule and check application-level identity where relevant. |
| A blocked client says robots.txt allowed the path | robots.txt is a crawler convention, not a grant of access. | Check the site’s access policy and use an authorized API or request permission. |
| One bot score is treated as definitive | A vendor-specific probabilistic score is being interpreted as proof. | Use the score only within its provider’s documented model and combine it with policy and context. |
| A managed feature is missing | The feature may depend on provider, configuration, or subscription plan. | Check the current product documentation and plan before building a rule around it. |
8. Performance, reliability, and cost considerations
Challenges add a step for users and can make automated API clients fail. Rate limits can protect costly operations but may also reject legitimate bursts. Machine-learning classifications and traffic baselines are probabilistic, so monitor false positives and false negatives rather than assuming detection is complete. Scope controls to the smallest useful set of routes and operations.
Operational cost includes configuring and reviewing rules, investigating reports, and maintaining exceptions for useful clients. Managed services differ in available signals and features by plan; compare current provider documentation and pricing for your deployment. The research sources do not provide a comparable independent benchmark for detection accuracy, latency, or cost, so those figures should not be inferred.
9. Frequently asked questions
Does a high request rate prove that traffic is scraping?
No. It can be one useful signal, but legitimate activity can also create bursts. Consider route, operation, identity, and context before taking action.
Can robots.txt stop a scraper?
No. It can guide compliant crawlers, but a noncompliant client can ignore it. Use server-side controls when access must be enforced.
Is Cloudflare’s bot score used by every provider?
No. The 1–99 range and the commonly associated below-30 range described here are specific to Cloudflare’s documented system.
Should every automated client be blocked?
No. Decide which automated traffic supports your site and which behavior causes harm, then apply policy accordingly.


