ScreenshotNeo

BlogGuides

Web Scraping Anti-Detection Techniques: A Practical Guide

Learn how sites identify automated traffic and how to collect public data responsibly with clear crawler identity, modest rates, caching, and permission.

By the ScreenshotNeo team4 October 202611 min read

There is no responsible way to guarantee that a scraper will avoid detection. The practical approach is to make authorized collection transparent and low impact: prefer an official data source, check the site’s rules, identify your crawler honestly, cache results, keep request rates modest, and stop when the site blocks access. A CAPTCHA, 403, or persistent rate limit is a signal to pause and ask for permission or use another supported route, not a challenge to bypass.

This guide explains what site operators may use to identify and limit automated traffic, then gives a runnable Python workflow for a permission-based crawler. It does not cover disguising automation, rotating residential proxies, spoofing browser fingerprints, or defeating CAPTCHAs.

1. What “anti-detection” should mean for an authorized crawler

In responsible crawling, the useful goal is not to conceal automation. It is to avoid behaving abusively or ambiguously: collect only what you are authorized to collect, make the client identifiable, reduce repeated work, respect published crawl guidance, and respond to overload or denial signals.

A publicly reachable URL is not automatically permission to collect it at scale. Terms, contracts, privacy obligations, copyright and database rules can depend on the site, jurisdiction, data, and purpose. This is operational guidance, not legal advice; get appropriate review when the collection is sensitive, regulated, or consequential.

2. How websites identify and limit automated traffic

Operators can assess request patterns and client-identification signals, and may use rate limits, CAPTCHA or other human verification, and bot mitigation. These controls can appear at the network, CDN, firewall, or application level. Their exact configuration varies, so a crawler should not infer that a page is fair to collect just because one request succeeds.

AWS describes client-identification controls and fingerprint-based rate limiting as bot-management approaches. OpenAI’s crawler guidance describes controls such as robots.txt, firewall or CDN protections, application-level verification, and throttling. These examples help explain why a request may be limited; they are not instructions for changing a client to evade those controls.

3. Check access rules before collecting

  1. Look for an official route first. Prefer an API, downloadable dataset, feed, or licensed source when one meets the need. Compare authorization clarity, freshness, coverage, rate limits, stability, cost, and privacy obligations.
  2. Read the site’s terms and crawler guidance. Check for a published crawler policy and the site’s robots.txt. Document the domains, paths, fields, purpose, and retention period in scope.
  3. Understand what robots.txt does. RFC 9309 specifies rules crawlers are requested to honor. It is crawler guidance, not an access-control mechanism or privacy wall; it does not guarantee that a page stays out of search results, and other crawlers may ignore it. Google’s documentation explains its interpretation for Google Search crawlers. [RFC 9309; Google Search Central: robots.txt]
  4. Get permission where needed. If the scope is unclear, ask the operator. Robots rules do not replace authorization or a legal assessment.

AWS’s ethical crawler guidance recommends checking site guidance, honoring robots.txt, and controlling crawl rate. Cloudflare’s sample terms are an example of one provider’s suggested wording, not a universal statement of law.

4. Build a transparent, low-impact workflow

  1. Identify the client truthfully. Use a descriptive user-agent that names your crawler and, where appropriate, gives a contact address or project page. Do not claim to be a search engine or another party’s client.
  2. Fetch only what the task needs. Keep a documented scope, avoid collecting unnecessary personal information, and set a retention limit.
  3. Use a conservative rate. Begin with one request at a time unless the operator has documented a higher allowance. Avoid request bursts and parallelism that could burden the site.
  4. Cache successful responses. Save results and avoid refetching unchanged pages. Respect cache directives and any site-specific rules; do not use caching to continue after an explicit denial.
  5. Handle transient failures with bounded backoff. A network error or server-side 5xx may be transient. Retry sparingly, with increasing delays and a strict cap. Treat 429, CAPTCHA, authentication barriers, 403, and persistent denial as stop signals.
  6. Keep an audit trail. Record the requested URL, time, status, retry count, and reason for stopping. Keep logs proportionate and avoid storing sensitive response content unnecessarily.

5. Runnable Python example for an authorized crawl

The example below uses Python’s standard library. It checks robots.txt for the crawler’s user-agent, fetches a single permitted page, identifies itself, caches the response on disk, and stops on rate limits, access denial, or human-verification pages. It intentionally does not retry a 403 or 429.

#!/usr/bin/env python3
"""Fetch one page for a permission-based, low-rate crawl."""
from pathlib import Path
from urllib.error import HTTPError, URLError
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
from urllib.request import Request, urlopen
import hashlib
import time

URL = "https://example.com/"
USER_AGENT = "ResearchCrawler/1.0 (+mailto:developer@example.org)"
CACHE_DIR = Path("crawl-cache")
MAX_BYTES = 2_000_000
TIMEOUT_SECONDS = 20


def cache_path(url: str) -> Path:
    key = hashlib.sha256(url.encode("utf-8")).hexdigest()
    return CACHE_DIR / (key + ".html")


def robots_allows(url: str) -> bool:
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    parser = RobotFileParser()
    parser.set_url(robots_url)
    try:
        request = Request(robots_url, headers={"User-Agent": USER_AGENT})
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            robots_text = response.read(MAX_BYTES).decode("utf-8", errors="replace")
        parser.parse(robots_text.splitlines())
    except HTTPError as exc:
        # A missing robots.txt is commonly treated as no published rules, but
        # an access or server error is not a reason to guess. Pause for review.
        if exc.code == 404:
            return True
        raise RuntimeError(f"Could not read robots.txt (HTTP {exc.code}); pause and check site guidance") from exc
    except (URLError, TimeoutError) as exc:
        raise RuntimeError("Could not read robots.txt; pause and check site guidance") from exc
    return parser.can_fetch(USER_AGENT, url)


def main() -> None:
    CACHE_DIR.mkdir(exist_ok=True)
    destination = cache_path(URL)
    if destination.exists():
        print(f"Cache hit: {destination}")
        return

    if not robots_allows(URL):
        raise SystemExit("robots.txt disallows this URL for this crawler")

    # One request at a time; the delay is an example courtesy pause, not a
    # claim that any particular rate is allowed by the target site.
    time.sleep(1)
    request = Request(URL, headers={"User-Agent": USER_AGENT, "Accept": "text/html"})
    try:
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            status = response.status
            content_type = response.headers.get("Content-Type", "")
            if status != 200:
                raise SystemExit(f"Unexpected HTTP status {status}; stop and review")
            if "text/html" not in content_type.lower():
                raise SystemExit(f"Expected HTML, received {content_type!r}")
            body = response.read(MAX_BYTES + 1)
            if len(body) > MAX_BYTES:
                raise SystemExit("Response exceeded the configured size cap; review before collecting")
            sample = body[:1000].lower()
            if b"captcha" in sample or b"verify you are human" in sample:
                raise SystemExit("Human-verification page detected; stop and ask the operator")
            if not body.strip():
                raise SystemExit("Empty response; stop and investigate rather than retrying blindly")
            destination.write_bytes(body)
            print(f"Saved {len(body)} bytes to {destination}")
    except HTTPError as exc:
        if exc.code in (403, 429):
            raise SystemExit(f"HTTP {exc.code}: access denied or rate limited; stop and seek an approved route") from exc
        if 500 <= exc.code < 600:
            raise SystemExit(f"HTTP {exc.code}: server error; pause and retry later only if site guidance permits") from exc
        raise SystemExit(f"HTTP {exc.code}: review the response and site rules") from exc
    except (URLError, TimeoutError) as exc:
        raise SystemExit(f"Network or timeout error: {exc}; pause and check connectivity and site guidance") from exc


if __name__ == "__main__":
    main()

Save it as crawl_one.py, replace URL and the contact in USER_AGENT, then run python3 crawl_one.py. The example is deliberately a one-page fetcher rather than a broad crawler: expanding it requires a reviewed URL scope, duplicate control, a conservative scheduler, and a clear stop policy.

6. Equivalent one-request examples in cURL and Node.js

These examples show how to identify a client and inspect a single authorized URL. They do not implement robots.txt checks or a crawl queue; add those checks before automating multiple URLs.

cURL

curl --fail --show-error --max-time 20 \
  -H 'User-Agent: ResearchCrawler/1.0 (+mailto:developer@example.org)' \
  -H 'Accept: text/html' \
  'https://example.com/' \
  -o page.html

Node.js

const url = 'https://example.com/';
const response = await fetch(url, {
  headers: {
    'User-Agent': 'ResearchCrawler/1.0 (+mailto:developer@example.org)',
    'Accept': 'text/html'
  },
  signal: AbortSignal.timeout(20_000)
});

if (response.status === 403 || response.status === 429) {
  throw new Error(`HTTP ${response.status}: stop and seek an approved access route`);
}
if (!response.ok) {
  throw new Error(`HTTP ${response.status}: review site rules and response`);
}
const contentType = response.headers.get('content-type') || '';
if (!contentType.toLowerCase().includes('text/html')) {
  throw new Error(`Expected HTML, received ${contentType}`);
}
const html = await response.text();
if (!html.trim()) throw new Error('Empty response; stop and investigate');
console.log(html.slice(0, 500));

7. Retry, caching, and crawl-rate decisions

Signal Responsible response
Network timeout or connection interruption Pause; check connectivity and whether the site is available. If retrying is within the site’s guidance, use a small bounded retry count with increasing delays.
HTTP 5xx Reduce activity and pause. Retry later only when the site’s guidance permits; stop if errors persist.
HTTP 429 Stop the crawl. Review any published rate guidance and contact the operator. Do not try to get around the limit.
HTTP 403, CAPTCHA, login, or human verification Stop. Ask for permission, use an official API or export, or abandon the collection.
Empty or unexpected content Do not assume the page is valid. Check content type, response status, and whether the site returned an error or verification page.
Repeated URLs or unchanged pages Use a cache and a canonical URL key to avoid duplicate fetches, subject to site rules.

There is no universal safe request rate. A site’s allowance depends on its own guidance and capacity. Start conservatively, avoid concurrent bursts, honor explicit limits, and ask the operator before increasing throughput. Cache controls repeated load and cost on your side, but does not grant permission to collect or override a site’s refusal.

8. Troubleshooting common crawler problems

Problem Likely cause What to do
Robots parser says the URL is disallowed A published rule applies to the crawler identity or path. Do not fetch that path. Check whether an official data route or explicit permission is available.
robots.txt cannot be fetched Network trouble, server error, or an access policy prevents reading it. Pause and consult the site’s published policy or operator; do not infer permission from a failed request.
HTTP 429 responses The service is rate limiting requests or signaling overload. Stop, review guidance, and ask the operator if the permitted rate is unclear. Do not keep retrying.
HTTP 403 or CAPTCHA The site denied access or requires human verification. Stop and use an official API, data export, license, or permission-based route.
HTTP 5xx or timeouts The site or network may be temporarily unavailable. Pause, check status through an approved channel, and only retry later if guidance permits. Bound retries.
Unexpected HTML or empty body The response may be an error, verification page, or a page that requires a supported access route. Inspect status and content type; do not escalate to evasion methods.
Too much duplicate data URLs may differ only by tracking parameters or the crawl may revisit pages. Define canonicalization rules for your dataset, deduplicate locally, and avoid requesting variants unless they are necessary and allowed.
Collection cost or storage grows unexpectedly The URL scope, response sizes, or retention period may be too broad. Limit fields and URLs, cap response sizes, cache results, and apply a documented retention policy.

9. Reliability, privacy, and cost notes

  • Reliability: a permission-based API or export is often easier to operate than parsing pages, when the source offers one. For a crawler, bounded timeouts, checkpointed progress, deduplication, conservative retries, and clear stop conditions reduce accidental load and make failures recoverable.
  • Performance: caching and avoiding duplicate requests improve throughput without sending more traffic. If greater speed is needed, ask the operator about an approved rate before using concurrency.
  • Cost: account for development time, network transfer, storage, maintenance when page structure changes, and any API or license fees. The research sources do not establish universal prices or a universally best collection route; compare actual options for the intended scope.
  • Privacy: collect only the fields the task requires, avoid sensitive personal data unless the use is approved and reviewed, limit access to stored results, and delete them when the purpose ends.

10. How site operators can reduce false positives

If you operate the site, bot controls should account for the cost of false positives, visitor friction, and operational effort. Make crawl rules and contact routes easy to find, publish reasonable rate expectations, and provide an official API or export where appropriate. OpenAI’s crawler guidance recommends that site operators review legitimate crawler access and investigate 429 responses in infrastructure logs; verified crawler traffic can be handled through documented policy. Avoid relying on opaque blocks when a clear contact and correction path would resolve legitimate access issues.

11. Or skip the browser setup

If the task is to capture a rendered screenshot of a page you are authorized to access, ScreenshotNeo provides a website screenshot API and MCP server. Its API takes one GET request and returns an image or PDF. This is a screenshot service, not a web scraper or a way to get around access controls: respect the target site’s rules and stop when it denies access.

One-call cURL example (replace the URL with a page you are authorized to capture):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python and Node.js requests, plus all available parameters, are in the ScreenshotNeo API documentation. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

12. FAQ

Does robots.txt grant permission to scrape?

No. It communicates crawler rules; it is not authorization, a privacy wall, or a complete statement of applicable law.

Is every automated request prohibited?

No. Some sites publish crawler guidance or offer APIs and data exports. Follow the site’s rules and get permission when required.

Should I retry a 429 or CAPTCHA?

No. Stop, review the published policy, and ask the operator or use a supported access route.

Can a screenshot API replace a crawler?

No. A screenshot API captures a rendered page as an image or PDF; structured data collection requires an authorized source and an appropriate extraction workflow.

Sources