ScreenshotNeo

BlogEngineering

Patterns and Anti-Patterns in Web Scraping

A practical guide to responsible scraping: robots.txt, resilient extraction, rate limits, browser automation, troubleshooting, and safer operations.

By the ScreenshotNeo team29 September 202610 min read

Patterns and Anti-Patterns in Web Scraping

Good web scraping starts with a narrow data requirement, a clear understanding of how the target serves content, and a client that responds to the target’s signals. The reliable pattern is to inspect the site, collect only the fields and pages you need, identify your crawler, use direct HTTP when the required data is in responses, use browser automation when rendered interaction is necessary, and slow down when the server asks you to.

The dangerous patterns are equally clear: treating robots.txt as a permission system, assuming one request rate works everywhere, retrying 429 responses in a tight loop, and tying extraction to fragile DOM structure. This guide explains the practical decisions, runnable implementations, failure handling, and operational checks behind a maintainable collector.

1. Start with a specific collection contract

Write down the exact URLs, fields, freshness requirement, and stopping condition before writing code. “Scrape the site” is not an actionable specification. A useful contract might say: collect the title, canonical URL, and price from product pages in a supplied list; run once daily; stop after three consecutive pages fail; retain the response status and extraction result for diagnosis.

  • Scope: list the hosts, paths, and page types you need.
  • Fields: define names, types, normalization, and what counts as missing.
  • Frequency: choose the slowest refresh that meets the business need.
  • Identity: send a descriptive product token in your User-Agent.
  • Audit trail: record URL, timestamp, status, response headers, parser version, and validation outcome.
  • Stop rules: decide when to pause for rate limits, repeated failures, or a changed layout.

Limit collection to relevant pages and fields. Data minimization is a practical engineering choice: it reduces load, storage, parsing work, and the amount of information that needs protection. It does not answer whether your project is legally or contractually permitted.

2. Read robots.txt correctly

robots.txt is crawler guidance, not authentication. RFC 9309 states that “These rules are not a form of access authorization.” An allowed path is not permission to access protected information, and a disallowed path is not a security barrier. Use authentication and authorization controls for sensitive resources.

Robots.txt guides crawlers but does not provide authorization.
Robots.txt guides crawlers but does not provide authorization.

Fetch the file for the exact host, protocol, and port you will request. Google’s documentation emphasizes that a robots file applies only to its own host, scheme, and port; do not apply rules from https://example.com to a different subdomain or HTTP endpoint. Match the user-agent group and the most specific applicable path rule as defined by RFC 9309.

from urllib.parse import urlparse
import urllib.robotparser


def robots_allows(url: str, user_agent: str = "AcmeCatalogBot/1.0 (+https://example.com/bot-info)") -> bool:
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    parser = urllib.robotparser.RobotFileParser(robots_url)
    parser.read()
    return parser.can_fetch(user_agent, url)

if not robots_allows("https://example.com/products/widget"):
    raise RuntimeError("Crawler policy does not allow this URL")

Production code should distinguish a successfully fetched, parseable policy from an unavailable or unreachable file. RFC 9309 provides standard guidance for those cases, while Google documents behavior specific to Google crawlers. Do not describe one implementation’s fallback as universal. Also avoid relying on a stale cached file indefinitely; RFC 9309 discusses cache freshness and a 24-hour limit unless the file is unreachable.

3. Choose direct HTTP or a browser

Question Direct HTTP client Browser automation
Where is the data? Use when the needed content is in the HTML or an API response. Use when user-visible rendering, JavaScript, scrolling, or interaction is required.
Operational shape Usually a smaller request and simpler parser, but still subject to server limits. More moving parts: browser version, scripts, waits, selectors, and resource loading.
Resilience Depends on response and markup stability. Prefer user-facing locators and explicit contracts; DOM-dependent selectors are fragile.
Throttling Honor status codes and Retry-After. Browser requests also reach the target and must honor the same signals.

This is a method-selection framework, not a speed or success benchmark. The cited sources do not establish universal performance differences.

4. A responsible direct-HTTP collector

The following Python example downloads a small, explicit URL list, identifies itself, applies a bounded delay, handles 429, and extracts a title. It deliberately does not implement an infinite retry loop.

import time
import requests
from bs4 import BeautifulSoup

USER_AGENT = "AcmeCatalogBot/1.0 (+https://example.com/bot-info)"


def fetch_title(url: str, session: requests.Session) -> dict:
    response = session.get(url, timeout=(10, 30), allow_redirects=True)
    result = {
        "url": url,
        "final_url": response.url,
        "status": response.status_code,
    }
    if response.status_code == 429:
        result["retry_after"] = response.headers.get("Retry-After")
        result["error"] = "rate_limited"
        return result
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    title = soup.title.get_text(" ", strip=True) if soup.title else None
    result["title"] = title
    result["has_title"] = bool(title)
    return result

urls = [
    "https://example.com/page-a",
    "https://example.com/page-b",
]

with requests.Session() as session:
    session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html"})
    for url in urls:
        try:
            print(fetch_title(url, session))
        except requests.RequestException as exc:
            print({"url": url, "error": str(exc)})
        time.sleep(2)

Use a queue and persistent result store for larger jobs. Keep concurrency low until you observe the target’s responses. A server’s policy, infrastructure, and rate limits vary, so no universal “safe” requests-per-second value can be inferred from the sources.

5. Handle status codes and retries

HTTP 429 Too Many Requests means the client sent too many requests in a period of time. MDN notes that the response may include Retry-After, which gives a wait duration or date. Treat a 429 as a signal to pause and reduce activity.

import random
import time


def retry_delay(response, attempt: int) -> float:
    retry_after = response.headers.get("Retry-After")
    if retry_after and retry_after.isdigit():
        return min(float(retry_after), 300.0)
    # Bounded exponential backoff with jitter; tune from observed policy.
    return min((2 ** attempt) + random.random(), 300.0)


def get_with_backoff(session, url, attempts=4):
    for attempt in range(attempts):
        response = session.get(url, timeout=30)
        if response.status_code != 429:
            return response
        if attempt == attempts - 1:
            return response
        time.sleep(retry_delay(response, attempt))
    raise AssertionError("unreachable")

Do not retry every error identically. A timeout may be transient; a consistent 404 usually is not. A 401 or 403 may indicate authentication, policy, or a blocked client and should trigger investigation rather than an aggressive retry loop. Record the response and stop conditions so operators can see whether the target changed.

6. Browser automation for rendered pages

Use Playwright when the value appears only after JavaScript execution, a click, a scroll, or another user-visible interaction. Playwright’s guidance favors user-facing locators and explicit contracts. Prefer a role, label, test id, or stable attribute over a selector that depends on a particular nesting structure.

Rendered interaction needs explicit waits and resilient locators.
Rendered interaction needs explicit waits and resilient locators.
import asyncio
from playwright.async_api import async_playwright


async def main():
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        page = await browser.new_page()
        await page.goto("https://example.com/catalog", wait_until="domcontentloaded")
        await page.get_by_role("button", name="Load more").click()
        await page.locator("[data-product-card]").first.wait_for()
        cards = await page.locator("[data-product-card]").all()
        for card in cards:
            name = await card.get_by_role("heading").inner_text()
            print(name)
        await browser.close()

asyncio.run(main())

Set explicit navigation and selector timeouts. Wait for the condition your extraction needs rather than using a large arbitrary sleep. If a page has infinite scrolling, define a maximum number of scrolls or items. Keep browser contexts isolated when cookies, locale, timezone, or authentication differ.

7. Content quality and change detection

A successful HTTP response is not necessarily a successful extraction. Validate required fields, canonical URLs, types, and reasonable lengths. Store a parser version and a small sample of raw input or a content hash where your retention policy permits it. A sudden rise in missing fields, identical pages, or bot-check text is an operational signal.

  • Check for an expected page marker before parsing.
  • Reject records missing required identifiers.
  • Compare field counts and distributions with recent runs.
  • Alert on layout changes instead of silently writing empty values.
  • Keep failed URLs for a bounded replay queue.

8. Common errors and fixes

Symptom Likely cause Fix
403 Forbidden Policy, authentication, client reputation, or an access rule. Review permission and terms, identify the client, reduce activity, and do not launch an immediate retry storm.
429 Too Many Requests Request rate exceeded a server threshold. Honor Retry-After, reduce concurrency, add bounded backoff, and resume slowly.
Empty HTML Content is rendered by JavaScript or a bot-check page was returned. Inspect the response; use the underlying documented endpoint if appropriate or browser automation when rendered output is required.
Selector timeout Wrong locator, changed markup, consent dialog, or incomplete navigation. Capture diagnostics, use a stable user-facing locator, wait for a specific condition, and handle dialogs.
Duplicate records Pagination overlap, redirects, retries, or unstable ordering. Deduplicate by a stable canonical key and record page boundaries.
Robots disagreement Rules belong to another host or crawler behavior was assumed from Google’s implementation. Fetch the exact host/scheme/port policy and distinguish RFC guidance from crawler-specific behavior.

9. Performance, reliability, and cost

Measure the parts you control: requests completed, bytes transferred, browser launches, queue age, extraction completeness, and retry counts. The research sources do not provide comparative benchmarks, universal safe rates, or cost figures, so avoid promising a fixed throughput.

For reliability, use connection pooling for direct HTTP, bounded concurrency, timeouts for connect and read phases, idempotent jobs, and durable checkpoints. For browser work, reuse a browser process while isolating contexts, block unnecessary resources only when it does not change the data contract, and pin compatible browser dependencies. Cache results when freshness allows it, but do not use a stale robots file as a substitute for current policy.

Cost includes bandwidth, compute, browser runtime, storage, and engineering time. Smaller scopes and slower refreshes reduce all of them. A managed screenshot endpoint can remove browser provisioning when your requirement is an image or PDF rather than structured fields.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

cURL (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

You can also use full-page capture with lazy images, CSS-element capture, device presets or custom viewports, dark mode, retina scale, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and PDF controls. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Plans include 1,000 free shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Technical behavior does not resolve permission. The cited materials do not establish jurisdiction-specific copyright, privacy, database-rights, contract, or reuse rules. Review the target’s terms, your agreements, the data subjects involved, retention requirements, and the intended downstream use. Escalate project-specific questions to qualified counsel rather than treating robots rules as a legal answer.

12. Practical checklist

  1. Define exact pages, fields, freshness, and stop conditions.
  2. Fetch and interpret the correct host’s robots policy for your crawler identity.
  3. Confirm contractual, privacy, legal, and reuse requirements.
  4. Choose direct HTTP for response-available data and browser automation for rendered interaction.
  5. Use descriptive identification, timeouts, bounded concurrency, and observable logs.
  6. Honor 429 and Retry-After; never run an immediate infinite retry loop.
  7. Prefer resilient, user-facing locators and explicit waits.
  8. Validate extracted fields and alert on layout or content changes.
  9. Retain only the data and logs your project needs.

FAQ

Does an allowed robots.txt path mean I may use the data?

No. RFC 9309 describes crawler guidance and explicitly says it is not access authorization. Permission and reuse depend on the target and project.

Should every scraper use Playwright?

No. Use a direct HTTP client when the required response is available without interaction. Use a browser when rendered output or interaction is part of the data contract.

How long should I wait after a 429?

Honor Retry-After when supplied. Otherwise pause with bounded backoff, reduce activity, and tune behavior from the target’s observed policy; there is no universal interval.

Can I solve a blocked page by changing selectors?

Usually not. A block, authentication requirement, or bot check is different from a selector failure. Inspect the response and resolve access and permission issues before changing extraction code.

When is a screenshot API a better fit?

When the output you need is a visual capture or PDF and you prefer a managed browser workflow. For structured fields, you still need an extraction method that validates the returned data.