ScreenshotNeo

BlogHow-to

How to Avoid Web Scraper Blocking

Learn how to crawl responsibly without triggering blocks: permission, robots.txt, APIs, honest identity, pacing, backoff, retries, and safer alternatives.

By the ScreenshotNeo team29 September 202610 min read

How to Avoid Web Scraper Blocking

Direct answer: The reliable way to avoid web scraper blocking is to make collection predictable and permitted. Check the site’s terms, robots.txt, authentication rules, and published limits; use an official API or export when one exists; identify your crawler honestly; keep concurrency and request rates conservative; cache results; and back off immediately on 429, 503, CAPTCHA, challenge, or ban responses. Never try to bypass an explicit denial. If you only need rendered page images, use a screenshot API instead of operating a browser fleet.

1. Start with permission, terms, and scope

Before writing code, define exactly what you will collect, from which host, how often, and for what purpose. Read the site’s terms of service, API documentation, authentication requirements, and any published automation policy. Look for a contact address or data-use policy if your project is commercial or high volume.

Robots.txt is part of the Robots Exclusion Protocol. RFC 9309 describes it as a request to crawlers and explicitly says, “These rules are not a form of access authorization.” In other words, a robots file helps you decide what to request, but it does not grant permission to access private data or override terms. Parse the file for your crawler’s user-agent group and honor the disallow rules that apply to your paths. A site can still deny access with authentication, a contract, a firewall, or an account-level policy.

Cache robots.txt rather than fetching it for every URL. RFC 9309 recommends a maximum cache period of 24 hours when the file is reachable. If the file cannot be fetched, record that condition and apply your organization’s policy; do not treat a network failure as permission to increase crawling.

2. Prefer an API, export, or search endpoint

An official API, bulk export, feed, or search endpoint usually reduces both your request count and the website’s work. Scrapy’s optimization guidance summarizes the reason clearly: “An API, a bulk export or a search endpoint is both faster for you and cheaper for the website than crawling its pages.” An API also gives you documented authentication, pagination, quotas, and error semantics.

Approach Best when Main trade-off
Official API You need structured, fresh records Quota, authentication, or endpoint cost
Bulk export You need a large historical or periodic dataset Updates may be delayed
Search/feed endpoint You need a narrow, discoverable subset Results may omit page details
HTML crawl No supported machine-readable source exists More requests, rendering complexity, and blocking risk

Use the smallest endpoint that answers your question. Do not crawl every product page if a catalog export contains the same fields. Do not fetch a full article repeatedly when an RSS or sitemap feed gives you change detection.

3. Identify your crawler honestly

Send a stable User-Agent that names the project and provides a contact or documentation URL where appropriate. Avoid impersonating a browser or rotating identities to conceal the same workload. RFC 9309’s matching model expects the product token in the User-Agent to correspond to the crawler’s identification string.

Mozilla/5.0 (compatible; CatalogResearchBot/1.0; +https://example.com/bot-info)

Keep the identity stable across workers and days. Log the User-Agent, source IP or egress pool, target host, request path, response status, and timing. If a site owner contacts you, pause the affected crawl and respond with your scope, rate, and contact details.

4. Pace requests and bound concurrency

Begin with one worker and a deliberate delay. Increase concurrency only after latency and error rates remain steady. A practical controller uses a per-host rate limit, a small maximum number of in-flight requests, and a queue that prevents duplicate URLs.

A permission check, queue, limiter, and cache keep collection predictable for the target site.
A permission check, queue, limiter, and cache keep collection predictable for the target site.

Translate the site’s published Crawl-delay or Request-rate into your downloader settings. Crawl during the target site’s local idle period when possible. The numeric examples in vendor documentation are illustrations, not universal limits: Cloudflare shows examples such as 10 requests in 2 minutes followed by 20 in 5 minutes, 50 in 10 seconds for a product lookup, and 5 in 1 hour for a GraphQL operation. Your safe rate depends on endpoint cost, traffic, identity, and the responses you observe.

import time
import threading

class HostRateLimiter:
    def __init__(self, min_interval_seconds=2.0):
        self.min_interval = min_interval_seconds
        self._lock = threading.Lock()
        self._next_allowed = 0.0

    def wait(self):
        with self._lock:
            now = time.monotonic()
            delay = max(0.0, self._next_allowed - now)
            self._next_allowed = max(now, self._next_allowed) + self.min_interval
        if delay:
            time.sleep(delay)

limiter = HostRateLimiter(2.0)
# Call limiter.wait() immediately before each request to this host.

Use separate limiters per hostname, and apply stricter limits to expensive paths such as search, login, GraphQL, and pages that trigger JavaScript rendering. Cache successful responses by URL and relevant headers. Normalize URLs and use a persistent deduplication store so restarts do not repeat work.

5. Handle 429, 503, challenges, and bans as stop signals

RFC 6585 defines HTTP 429 as “Too Many Requests” and says a response may include Retry-After. Parse both the delta-seconds and HTTP-date forms. For 503, CAPTCHA pages, bot challenges, or a clear ban page, stop or sharply reduce traffic; do not respond by adding more parallel workers or rotating proxies.

import email.utils
import random
import time
from datetime import datetime, timezone


def retry_after_seconds(value, default=60):
    if not value:
        return default
    try:
        return max(0, int(value))
    except ValueError:
        try:
            when = email.utils.parsedate_to_datetime(value)
            if when.tzinfo is None:
                when = when.replace(tzinfo=timezone.utc)
            return max(0, int((when - datetime.now(timezone.utc)).total_seconds()))
        except (TypeError, ValueError, OverflowError):
            return default


def backoff(attempt, retry_after=None):
    if retry_after is not None:
        return retry_after
    # Exponential backoff with bounded jitter.
    return min(900, (2 ** attempt) + random.uniform(0, 1))

# On 429: sleep for Retry-After, then retry only a bounded number of times.
# On repeated 429/503 or a challenge page: stop the host queue and review policy.

Retry only transient failures. A timeout may be retried with backoff; a 401, 403, robots denial, or explicit account suspension requires a policy decision, not blind retries. Set a maximum attempt count and persist failed URLs for review.

6. A conservative Python crawler skeleton

The following example uses only the standard library plus requests. It checks robots.txt, identifies itself, limits one host to one request every two seconds, honors Retry-After, and stops on common block signals. Replace the example domain and paths only after confirming permission.

import time
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser

import requests

START_URLS = ["https://example.com/catalog"]
USER_AGENT = "CatalogResearchBot/1.0 (+https://example.com/bot-info)"
MIN_INTERVAL = 2.0
TIMEOUT = 30

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
last_request = {}
robots = {}


def allowed(url):
    parts = urlparse(url)
    origin = f"{parts.scheme}://{parts.netloc}"
    if origin not in robots:
        rp = RobotFileParser(urljoin(origin, "/robots.txt"))
        try:
            rp.read()
        except OSError:
            return False  # Apply a fail-closed policy when robots cannot be read.
        robots[origin] = rp
    return robots[origin].can_fetch(USER_AGENT, url)


def fetch(url, attempts=3):
    if not allowed(url):
        raise RuntimeError(f"robots.txt disallows or is unavailable: {url}")
    host = urlparse(url).netloc
    for attempt in range(attempts):
        elapsed = time.monotonic() - last_request.get(host, 0)
        if elapsed < MIN_INTERVAL:
            time.sleep(MIN_INTERVAL - elapsed)
        response = session.get(url, timeout=TIMEOUT)
        last_request[host] = time.monotonic()
        if response.status_code == 200:
            body_start = response.text[:2000].lower()
            if any(marker in body_start for marker in ("captcha", "verify you are human", "access denied")):
                raise RuntimeError("challenge or ban page detected; stopping")
            return response
        if response.status_code in (429, 503):
            wait = retry_after_seconds(response.headers.get("Retry-After"), 60)
            time.sleep(wait)
            continue
        if response.status_code in (401, 403):
            raise RuntimeError(f"access denied with HTTP {response.status_code}")
        response.raise_for_status()
    raise RuntimeError("retry limit reached")


def retry_after_seconds(value, default):
    try:
        return max(0, int(value)) if value else default
    except ValueError:
        return default

for url in START_URLS:
    response = fetch(url)
    print(response.url, len(response.content))
    # Parse only the fields you need, enqueue only in-scope links, and deduplicate.

For production, add durable queues, per-host configuration, HTML parsing limits, content hashing, metrics, and an operator-controlled stop switch. Keep a record of why each URL was skipped.

7. JavaScript and authentication edge cases

Client-rendered pages can make a simple HTTP request look empty. First check whether the data is available through an API call visible in the site’s documentation or network panel. If authentication is required, obtain credentials through the documented process and protect cookies and tokens. Do not defeat login controls, CAPTCHA, paywalls, or access checks.

Some pages vary by timezone, geolocation, cookies, or user-agent. Record those inputs with the fetched artifact so results are reproducible. Avoid sending unnecessary cookies or personal data. If a page requires a browser only to render a public view, a screenshot service can isolate that browser complexity from your crawler.

8. Or skip the browser setup

If your goal is a visual capture rather than structured extraction, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP, or PDF. It handles the browser session and includes controls for full-page capture, lazy images, CSS selectors, dark mode, device presets, custom viewport and retina scale, waits, custom CSS and JavaScript, clicks, hidden selectors, blocked resource types, headers, cookies, user-agent, authorization, timezone, geolocation, transparency, resizing, caching, signed links, async jobs, webhooks, bulk capture, and usage reporting. See the ScreenshotNeo API documentation for parameter names and response details.

A rendering workflow can remove obstructive overlays before producing the final capture.
A rendering workflow can remove obstructive overlays before producing the final capture.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. Troubleshooting common blocks

Symptom Likely cause Fix
429 with Retry-After Rate limit exceeded Honor the header, lower concurrency, and add per-host pacing.
403 or access denied Policy, WAF, authentication, or ban Stop, verify permission, and contact the owner. Do not rotate identities to evade it.
503 spikes Server overload or protective control Pause the queue, back off, and schedule during idle hours.
CAPTCHA or challenge HTML Bot detection triggered Stop automated access and use an approved API or request authorization.
Empty HTML Content rendered by JavaScript Find the documented data endpoint or use an authorized rendering workflow.
Duplicate downloads No durable URL normalization or cache Canonicalize URLs, hash responses, and persist crawl state.
Robots parser says disallowed User-agent group or path rule matches Skip the URL and record the rule; ask the owner if you need access.

10. Performance, reliability, and cost

  • Measure: Track requests per minute, status counts, p50/p95 latency, bytes, retries, and challenge detections by host.
  • Reduce work: Use conditional requests such as ETag or Last-Modified when supported, cache pages, and crawl only changed URLs.
  • Protect reliability: Use bounded queues, timeouts, circuit breakers, and resumable checkpoints. A slow crawl that finishes is more useful than a fast crawl that gets banned.
  • Control spend: APIs and exports often reduce bandwidth and parsing costs. For screenshots, caching and bulk capture can reduce repeated browser work; ScreenshotNeo lets you choose a cache TTL and supports up to 100 URLs per bulk call.
  • Plan capacity: Treat vendor examples such as 50 requests per 10 seconds as policy illustrations, not targets. Derive limits from the site’s documentation and observed responses.

11. Practical preflight checklist

  1. Confirm the target permits the intended collection.
  2. Read terms, authentication rules, API limits, and robots.txt.
  3. Choose an API, export, or feed when available.
  4. Set an honest, stable User-Agent with contact information.
  5. Configure conservative per-host delay and bounded concurrency.
  6. Cache responses and deduplicate URLs.
  7. Detect 429, 503, CAPTCHA, challenge, and ban pages.
  8. Honor Retry-After and stop when access is denied.
  9. Keep logs, checkpoints, and a manual stop switch.
  10. Contact the site owner before requesting a higher limit.

12. FAQ

Does robots.txt stop scraping?

No. It communicates crawler preferences; RFC 9309 says it is not access authorization. You should honor it and separately comply with terms, authentication, and direct denials.

How long should I wait after a 429?

Use Retry-After when supplied. Otherwise apply exponential backoff with jitter, reduce concurrency, and stop after a bounded number of attempts.

Should I use proxies to avoid blocking?

Do not use proxy rotation to evade a site’s limits or ban. If you have authorized distributed collection, document the approved egress and still honor the site’s rate policy.

What is a safe crawl speed?

There is no universal number. Start slowly, follow published limits, watch latency and errors, and increase only with evidence that the host can handle it.

When is a screenshot API better than scraping?

Use one when you need a rendered image or PDF, especially for JavaScript-heavy pages, and do not need the page’s structured data. It avoids maintaining your own browser fleet while preserving a controlled request rate.