ScreenshotNeo

BlogGuides

How Websites Detect and Prevent Web Scraping

Websites detect scraping through layered signals, then choose whether to monitor, rate-limit, challenge, or block. Learn what works, what does not, and how to tune defenses safely.

By the ScreenshotNeo team4 October 20269 min read

Websites detect likely scraping by combining request details, known bot signatures, browser checks, behavioral signals, and traffic patterns. They can respond by logging, rate-limiting, challenging, or blocking requests. No single signal proves that a visitor is scraping, and robots.txt is a crawler preference file—not access control. Protect private data with authentication and authorization.

This guide is for developers and site operators designing defenses. It explains the signals, response choices, rollout process, common mistakes, and operational tradeoffs. It does not provide a universal threshold: traffic, endpoint cost, user expectations, and legitimate automated clients differ by site.

1. How websites detect scraping

Detection is a classification problem. A site or its web application firewall (WAF) collects signals and estimates whether a request or session is automated. Operators then decide what action, if any, is appropriate. Vendor documentation describes available methods; it is not independent evidence of any product’s detection accuracy.

Signal family What it can reveal Limit
Request attributes User-agent, request headers, IP reputation, request rate, and URL or parameter patterns can identify obvious automation or known clients. Headers can be copied or changed. Shared networks and legitimate scripts can resemble automation.
Known bot identification Bot controls can classify self-identifying crawlers and, for some known crawlers, verify that they originate from the organization they claim to represent. This primarily addresses known or self-identifying bots. It does not establish that every unclassified request is human.
Browser and connection checks Browser interrogation and TLS fingerprints can contribute evidence about the client making a request. These are signals, not identity or intent proofs. Client changes and intermediaries complicate interpretation.
Behavior Navigation patterns, timing, and session context can help distinguish ordinary browsing from repeated extraction. Accessibility tools, monitoring, and unusual but legitimate users can have atypical behavior.
Aggregate traffic analysis Patterns across requests, clients, networks, or fingerprints can reveal coordinated activity that a single request does not show. Requires context and tuning. A shared network or popular client characteristic is not itself evidence of abuse.

AWS documents common protections for self-identifying bots and targeted protections that use browser interrogation, TLS fingerprints, behavioral heuristics, and traffic analysis. Cloudflare documents scraping detections that analyze request patterns by ASN and JA4 fingerprint and dynamically recalculate matches. Treat these as descriptions of vendor capabilities, not comparative performance claims. AWS WAF Bot Control, AWS Bot Control use cases, and Cloudflare scraping detections.

The practical rule is to combine context. A high request rate might be a crawler, a mobile application, a monitoring job, or a busy shared NAT address. A suspicious user-agent alone can be forged. Use multiple signals and endpoint context, then select a proportional response.

2. Choose a response that matches the evidence

Detection and mitigation are separate decisions. A useful response ladder starts with the least disruptive action that protects the operation:

  1. Observe: log classifications, endpoint, request key, and outcome. Avoid collecting more personal data than needed.
  2. Classify: distinguish known search crawlers, partner integrations, your own jobs, unknown automation, and clearly abusive traffic.
  3. Limit: slow or cap expensive operations, ideally with a response that clients can handle.
  4. Challenge: ask uncertain browser sessions to demonstrate browser capability, or use CAPTCHA where appropriate.
  5. Block: deny traffic when evidence, policy, and impact review justify the action.

AWS WAF describes a silent Challenge that checks whether a session is a browser, and CAPTCHA as an interactive puzzle. Challenges can reduce the impact on legitimate requests compared with immediate blocking, but they add friction and may not suit APIs or non-browser clients. AWS also documents additional costs for Bot Control and CAPTCHA or Challenge actions; check current service terms and pricing before deployment. AWS WAF CAPTCHA and Challenge.

3. Scope rate limits to the operation

Do not assume one request threshold fits every route. Apply controls to high-cost or high-value operations—such as price lookups, search, or bulk export—and choose a key that matches how the application identifies legitimate clients.

  • IP address: useful for simple abuse patterns, but shared networks can group unrelated users and distributed clients can evade per-IP limits.
  • Session cookie: useful for browser flows where the application has a meaningful session. It is unsuitable as the only identity for unauthenticated or non-browser API clients.
  • Account or API credential: can align limits with a customer or integration, provided credentials are handled safely and unauthenticated routes have their own controls.
  • Operation and parameters: separate costly query variants or endpoints so a cheap page view does not share a budget with an expensive lookup.

Cloudflare’s rate-limiting guidance gives examples keyed by IP, query parameters, or a session cookie, with actions such as challenge and block. Its example values are illustrations, not recommended universal thresholds. Start from your own legitimate traffic distribution, capacity, and abuse impact. Cloudflare rate-limiting best practices.

4. Roll out defenses without blocking good traffic

  1. Inventory protected routes. Identify public content, expensive operations, authenticated data, APIs, and partner integrations. Rank routes by cost and impact.
  2. Establish a baseline. Review ordinary traffic by route, client type, time, and relevant session or account key. Include scheduled jobs and legitimate crawlers.
  3. Enable monitor or count mode. Record what a proposed rule would match without enforcing it. AWS recommends count-mode review before switching Bot Control rules to blocking.
  4. Inspect false positives. Check affected users, API clients, search crawling, mobile clients, and monitoring. AWS recommends reviewing labels and using application SDK signals when evaluating targeted protection that depends on client-side session context.
  5. Enforce narrowly. Start with the routes and classifications where the evidence is strongest. Exclude or adapt API calls that cannot complete browser challenges.
  6. Watch outcomes and tune. Track allowed, challenged, throttled, and blocked traffic; errors, support reports, and capacity. Keep a rollback path and revisit rules as traffic and product behavior change.

AWS describes count mode followed by review of labels and false positives as a deployment approach. AWS: choosing and configuring Bot Control.

5. What robots.txt does—and does not—do

robots.txt communicates crawler preferences. Compliant crawlers may use it to decide what to fetch; a noncompliant crawler can ignore it. It does not authenticate a visitor, authorize access, or protect confidential files. Google cautions that it should not be used to hide pages from Search: a disallowed URL can still appear in results if other pages link to it. Google’s robots.txt guide.

IETF RFC 9309 makes the security boundary explicit: “The Robots Exclusion Protocol is not a substitute for valid content security measures” and its rules “are not a form of access authorization.” Put private data behind actual access controls, such as authentication and authorization; Google recommends password protection for private files. RFC 9309.

6. Protect data with access control

If a response must be private, make the application check who is requesting it and whether that identity is allowed to access the specific resource. Do not rely on an obscure URL, a disallow rule, a browser challenge, or a bot score. Those measures can reduce unwanted crawling, but they are not substitutes for authorization. Apply least privilege, avoid exposing sensitive data in public endpoints, and review access logs for unexpected use.

7. Inspect pages with a screenshot API

When you need to review how a public page renders, a screenshot can help identify consent overlays, blank states, or layout changes without building a browser capture pipeline. ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. It returns PNG, JPEG, WebP, or PDF from a GET request. Its consent handling accepts cookie banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. This is for page rendering review, not a way to bypass a site’s access controls. See ScreenshotNeo and the API documentation.

8. Or skip the browser setup

For an authorized page, one request returns the capture. Replace YOUR_API_KEY with your key and change the target URL as needed. The endpoint and parameters are documented at ScreenshotNeo docs.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed; response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.

9. Troubleshooting common defense problems

Symptom Likely cause What to do
Legitimate users are blocked A broad rule, shared IP, or single weak signal is being treated as proof. Return to count mode, inspect matched routes and client context, narrow the rule, and allow known integrations where justified.
Scraping continues after an IP limit Clients may be distributed across addresses, or the protected route/key may not match the actual operation. Review traffic across sessions and endpoints; choose a key aligned with the application identity and operation. Combine signals rather than assuming IP is sufficient.
A browser challenge breaks an API client The client cannot execute browser checks or solve an interactive challenge. Scope challenges to browser routes, provide a suitable authenticated API path, or exclude the API call from that action.
robots.txt disallow has no effect The crawler may not follow the preference file. Use robots.txt only for compliant crawler guidance. Enforce access control for private material and use WAF or application controls for abuse mitigation.
Known crawlers are classified incorrectly A self-claimed user-agent may not match a verified crawler, or legitimate bot categories have not been allowed appropriately. Use the provider’s verification/classification signals and review the rule outcome before enforcement.
Rules work in logs but cause unexpected cost or friction Managed inspection or challenge actions can have service charges and user impact. Check current provider pricing and action requirements; measure challenge rates and legitimate completion before expanding enforcement.

10. Performance, reliability, and cost

  • Performance: rate-limit the expensive operation to protect capacity, and avoid putting unnecessary challenge work on every request. The dossier does not establish comparable latency figures, so measure your own routes and configuration.
  • Reliability: count-mode rollout, visible classifications, narrow rules, and a rollback path reduce the risk of an accidental outage. Review false positives after application or traffic changes.
  • Cost: managed bot inspection and challenge actions may carry extra provider fees. Estimate based on the provider’s current pricing and expected action volume; no cross-vendor price or effectiveness benchmark is established here.
  • Operations: advanced targeted protections may need client-side SDK signals or integration. Account for integration, monitoring, and tuning effort when selecting controls.
  • User experience: CAPTCHA adds work for people. Prefer a proportionate response and consider API clients and assistive or unusual browsing patterns when tuning.

11. Frequently asked questions

Can websites tell if you are scraping?

They can classify traffic as likely automated using combined signals, but that classification is contextual and not certainty about intent. An individual request attribute is not proof.

Does robots.txt stop scraping?

No. It expresses preferences for compliant crawlers. It is not an enforceable barrier or authorization mechanism.

How do websites block bots?

Operators can combine bot classification, scoped rate limits, browser challenges, CAPTCHA, and blocking rules. They should monitor first and review false positives before enforcement.

Should every automated client be blocked?

No. Search crawlers, monitoring, partner integrations, and application clients may be legitimate. Classify and apply policy by route and client purpose.

Can a challenge protect private data?

No. A challenge can help manage suspicious sessions, but private data needs authentication and authorization at the application or service boundary.