ScreenshotNeo

BlogGuides

13 Tips to Master Data Crawling: Building Reliable Crawls

Build reliable web crawls with 13 practical tips for scope, robots.txt, pacing, retries, caching, monitoring, extraction quality, and provenance.

By the ScreenshotNeo team1 October 202611 min read

Short answer: reliable data crawling starts with a narrow data question, an intentional URL inventory, permission-aware access, conservative per-host pacing, adaptive backoff, caching, resilient extraction, and complete observability. The 13 tips below turn those principles into an operating checklist you can use for a small script or a long-running crawl.

This guide covers independent crawlers. Google’s crawl-budget guidance describes Google’s systems and site-owner controls; it is useful diagnostic context, not a universal rule for every crawler. AWS’s request-rate examples are also context-specific. Always follow the destination’s instructions and adapt to its responses.

1. Define the data question and fields first

Write down the decision your dataset must support before discovering URLs. Specify the fields, types, acceptable missingness, freshness target, and what counts as a valid record.

  • Scope: domains, subdomains, paths, languages, and time range.
  • Fields: names, types, normalization rules, and required versus optional values.
  • Freshness: one-time snapshot, daily update, or change-driven recrawl.
  • Quality rule: for example, reject a product record unless it has a name and canonical URL.
  • Provenance: retain source URL, retrieval time, response status, parser version, and content hash.

A written schema prevents a crawler from collecting thousands of pages that cannot answer the original question.

2. Check for an API or bulk dataset

Before crawling HTML, look for a documented API, export, feed, or bulk download. A maintained interface usually gives you stable fields, explicit limits, and less load on the site. The W3C Data on the Web Best Practices recommend standards-based APIs, complete documentation, and communication of breaking changes: W3C Data on the Web Best Practices.

Compare routes using permission, request burden, useful coverage, freshness, resilience to page changes, data quality, provenance, and operating cost. If an API is incomplete, use it for the fields it provides and crawl only the documented gaps.

3. Read robots.txt and access requirements

Fetch and review /robots.txt before scheduling URLs. Treat it as a publisher’s crawler preference signal, not as an access-control mechanism for confidential data. Never access private or login-protected content without authorization.

from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser


def allowed(url: str, user_agent: str) -> bool:
    parts = urlparse(url)
    robots_url = f'{parts.scheme}://{parts.netloc}/robots.txt'
    parser = RobotFileParser(robots_url)
    parser.read()
    return parser.can_fetch(user_agent, url)

print(allowed('https://example.com/catalog/item-1', 'ExampleResearchBot/1.0 (+https://example.org/contact)'))

Cache robots.txt for a reasonable period and refresh it during long jobs. If the file is unavailable, choose a conservative policy, document it, and contact the site owner when the project is material.

4. Identify your crawler clearly

Use a descriptive user-agent that names your project and provides a contact URL or email where appropriate. Do not impersonate a browser or another crawler. A clear identity lets operators report problems and helps you distinguish your traffic in logs.

USER_AGENT = 'ExampleResearchBot/1.0 (+https://example.org/crawler-info)'
HEADERS = {'User-Agent': USER_AGENT, 'Accept': 'text/html,application/xhtml+xml'}

Keep the identity stable across runs so rate limits, allowlists, and incident analysis remain meaningful.

Use sitemap files, documented feeds, and crawlable links to build the inventory. Sitemaps are hints about important or recently changed URLs; they do not guarantee immediate fetching. Parse sitemap indexes, normalize URLs, and record where each URL was discovered.

import requests
from xml.etree import ElementTree as ET


def sitemap_urls(sitemap_url):
    xml = requests.get(sitemap_url, headers=HEADERS, timeout=30)
    xml.raise_for_status()
    root = ET.fromstring(xml.content)
    return [node.text.strip() for node in root.findall('.//{*}loc') if node.text]

for url in sitemap_urls('https://example.com/sitemap.xml'):
    print(url)

Extract links from HTML only within your allowed host and path boundaries. Store discovery source, canonical URL, and first-seen time for auditability.

6. Bound the URL space and remove duplicates

URL spaces can expand without limit through query parameters, calendars, filters, sessions, and tracking links. Define an allowlist and canonicalization policy before crawling.

  • Restrict hosts, schemes, ports, and path prefixes.
  • Drop known tracking parameters such as campaign identifiers when they do not change content.
  • Sort query parameters only when the site treats their order as irrelevant.
  • Deduplicate fragments because they usually select a position within the same document.
  • Set maximum depth, page count, response size, and runtime.
  • Use a content hash to detect duplicate responses served at different URLs.

Keep rejected URLs with a reason. This makes coverage decisions explainable instead of silently losing data.

7. Set a conservative per-host pace

Rate-limit independently per host, not just globally. AWS gives examples of one request every 10–15 seconds for small or medium sites and one to two requests per second for larger sites or explicit permissions; these are contextual examples, not blanket safe limits. Start slowly, observe server behavior, and increase only with permission and evidence.

import random
import time

MIN_DELAY = 10.0
MAX_DELAY = 15.0


def wait_between_requests():
    time.sleep(random.uniform(MIN_DELAY, MAX_DELAY))

For multiple workers, use a shared host-level token bucket so concurrency cannot bypass the delay. Batch long jobs into windows and avoid synchronized bursts at the top of an hour.

8. Back off on 429, 5xx, and slow responses

Slow responses, overload signals, HTTP 429, and repeated 5xx responses mean traffic should decrease. AWS recommends pausing on 429 and considering a stop if 403 responses persist. Google’s crawl-capacity guidance likewise describes slower responses, 5xx errors, and 429s as signals that reduce crawl capacity.

import time
import requests

RETRYABLE = {429, 500, 502, 503, 504}


def get_with_backoff(url, session, attempts=5):
    for attempt in range(attempts):
        response = session.get(url, headers=HEADERS, timeout=(10, 60))
        if response.status_code not in RETRYABLE:
            return response
        retry_after = response.headers.get('Retry-After')
        delay = float(retry_after) if retry_after and retry_after.isdigit() else min(300, 2 ** attempt * 10)
        time.sleep(delay)
    raise RuntimeError(f'Giving up after repeated transient errors: {url}')

Do not retry every failure. A persistent 403 requires investigation and usually a stop for that host. A 404 is normally terminal. Record status, attempt count, delay, and the final decision.

9. Cache unchanged content and use conditional requests

Cache successful responses when freshness requirements allow it. Preserve ETag and Last-Modified values, then send If-None-Match or If-Modified-Since on later runs. A 304 response lets you reuse the stored body without downloading it again. Google lists HTTP caching and 304 responses as ways to save bandwidth.

def conditional_get(url, session, metadata):
    headers = dict(HEADERS)
    if metadata.get('etag'):
        headers['If-None-Match'] = metadata['etag']
    if metadata.get('last_modified'):
        headers['If-Modified-Since'] = metadata['last_modified']
    return session.get(url, headers=headers, timeout=(10, 60))

Use a content-addressed store or database keyed by canonical URL and retrieval version. Never let a stale cache silently satisfy a run whose freshness contract requires current data.

10. Handle redirects and terminal statuses deliberately

Follow a limited number of redirects and record the complete chain. Long chains waste time and can hide loops. Store the final URL, status, and redirect targets so canonicalization rules can be corrected.

  • 2xx: parse only after checking content type and size.
  • 3xx: follow within policy; stop on loops or excessive hops.
  • 401/403: do not bypass; verify authorization and access instructions.
  • 404/410: mark terminal and remove from active queues unless the project explicitly rechecks removals.
  • 429: honor Retry-After and reduce host traffic.
  • 5xx: retry with backoff, then quarantine the URL or host.

11. Make extraction resilient to page changes

Selectors break when templates change, content moves behind JavaScript, or a site serves different variants. Prefer semantic signals and stable attributes over deeply nested CSS paths. Validate every record before writing it.

from bs4 import BeautifulSoup
from urllib.parse import urljoin


def extract_record(html, page_url):
    soup = BeautifulSoup(html, 'html.parser')
    title = soup.select_one('h1, [data-testid="title"]')
    canonical = soup.select_one('link[rel="canonical"]')
    record = {
        'url': page_url,
        'canonical_url': urljoin(page_url, canonical.get('href')) if canonical else page_url,
        'title': title.get_text(' ', strip=True) if title else None,
    }
    if not record['title']:
        raise ValueError('missing required title')
    return record

Keep parser versions with outputs. For pages that require rendering, separate browser acquisition from extraction, set a finite wait condition, and capture diagnostics such as final URL, response status, and a small HTML sample when permitted.

12. Monitor requests, coverage, and data quality

A crawl is reliable only when you can explain what happened. Emit structured events for every URL and aggregate them by host, status, attempt, latency, bytes, parser result, and reason for exclusion.

Track at least:

  • discovered, scheduled, fetched, parsed, rejected, and retried counts;
  • status-code and content-type distributions;
  • latency percentiles, timeout rate, and bytes transferred;
  • robots exclusions, redirect depth, and duplicate rate;
  • required-field completeness and validation failures;
  • coverage by path, sitemap, language, and discovery source.

Keep discovery, access, and downstream use separate in reports. Google’s troubleshooting guidance explicitly reminds site owners to distinguish crawling from indexing; your own dataset should make the same separation.

13. Preserve provenance, versions, and change history

Store enough metadata to reproduce and audit each record:

  • source URL and final URL;
  • retrieval timestamp and timezone;
  • HTTP status, selected headers, and content hash;
  • crawler and parser version;
  • robots decision and user-agent;
  • discovery source and crawl run identifier;
  • validation outcome and transformation steps.

Keep raw responses only when your legal, privacy, and retention policies permit it. Otherwise retain hashes, extracted fields, and a traceable reference to the source. Version schemas and parsers so a later run can be compared with an earlier one.

Reference implementation: a small respectful crawler

The following Python example combines bounded discovery, robots checks, per-host pacing, retries, and basic extraction. It is intentionally small; production crawlers should add durable queues, shared rate limiting, persistent caching, and metrics.

import time
import requests
from collections import deque
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
from bs4 import BeautifulSoup

START = 'https://example.com/'
USER_AGENT = 'ExampleResearchBot/1.0 (+https://example.org/crawler-info)'
ALLOWED_HOST = urlparse(START).netloc
MAX_PAGES = 100
DELAY = 10

session = requests.Session()
session.headers.update({'User-Agent': USER_AGENT, 'Accept': 'text/html,application/xhtml+xml'})
robots = RobotFileParser(urljoin(START, '/robots.txt'))
robots.read()
queue = deque([START])
seen = {START}

while queue and len(seen) <= MAX_PAGES:
    url = queue.popleft()
    if not robots.can_fetch(USER_AGENT, url):
        continue
    time.sleep(DELAY)
    try:
        response = session.get(url, timeout=(10, 60), allow_redirects=True)
    except requests.RequestException:
        continue
    if response.status_code == 429:
        time.sleep(60)
        continue
    if response.status_code != 200 or 'text/html' not in response.headers.get('content-type', ''):
        continue
    soup = BeautifulSoup(response.text, 'html.parser')
    title = soup.select_one('h1')
    if title:
        print({'url': response.url, 'title': title.get_text(' ', strip=True)})
    for link in soup.select('a[href]'):
        target = urljoin(response.url, link['href']).split('#', 1)[0]
        parsed = urlparse(target)
        if parsed.scheme in ('http', 'https') and parsed.netloc == ALLOWED_HOST and target not in seen:
            seen.add(target)
            queue.append(target)

Performance, reliability, and cost planning

Throughput

Throughput is constrained by the slowest destination and your allowed request rate. More workers do not make a crawl faster when the host-level limiter is the bottleneck. Use bounded concurrency across hosts while preserving a separate conservative budget for each host.

Reliability

Use durable queues so a process restart does not lose URLs. Make writes idempotent with a stable record key. Quarantine repeatedly failing hosts and URLs instead of retrying forever. Keep a dead-letter list for manual review.

Cost

Measure requests, transferred bytes, browser-rendered pages, storage, and processing time. APIs and bulk downloads can reduce request volume. Conditional requests, caching, deduplication, and early rejection of out-of-scope URLs reduce both network and compute cost. For large workloads, size capacity from measured latency and response size rather than assuming a universal crawl rate.

Or skip the browser setup

If your crawl needs screenshots for visual verification, rendered-page evidence, or a quick check that a page loaded correctly, ScreenshotNeo provides a single HTTP endpoint. It can load lazy images, capture full pages or one CSS-selected element, apply custom waits, headers, cookies, user agents, timezone and geolocation settings, and return PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

cURL:

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);

See the ScreenshotNeo API documentation for all options. ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots each month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.

Troubleshooting checklist

Symptom Likely cause Fix
Many 429 responses Per-host rate is too high or requests arrive in bursts Honor Retry-After, lower concurrency, add jitter, and pause the host.
Persistent 403 responses Access policy, authorization, or blocked identity Stop, review robots and terms, contact the owner, and do not bypass controls.
Queue grows forever Unbounded parameters, calendars, or session URLs Apply host/path allowlists, canonicalization, depth and page limits.
Records suddenly lose fields Template or selector change Fail validation, retain samples, version the parser, and update selectors.
High duplicate rate Tracking parameters or alternate URL forms Normalize URLs and compare content hashes before storing records.
Frequent timeouts Slow origin, oversized pages, or an unsuitable timeout Reduce rate, set connect/read timeouts separately, and quarantine repeated failures.
Freshness is unclear Cache has no explicit policy Store retrieval time, validators, TTL, and the freshness target for each dataset.
Results cannot be audited Missing source and parser metadata Persist URL, final URL, status, hash, run ID, and parser version with every record.

FAQ

Does robots.txt grant permission to crawl?

No. It communicates crawler preferences. It does not authorize access to private, restricted, or confidential data.

How fast should a crawler send requests?

There is no universal safe rate. Follow the site’s instructions, begin conservatively, and reduce traffic when latency or error rates rise. AWS’s published examples are context-specific guidance.

Should every sitemap URL be crawled?

No. Treat sitemaps as discovery and freshness hints. Apply your scope, duplicate, value, and freshness rules.

Is a 304 response a failed crawl?

No. It means the representation has not changed according to the server’s validator; reuse the cached body and record the validation event.

Does crawling mean the data is complete?

No. Measure coverage against your defined inventory and report exclusions, failures, and unknowns. A fetched page can still fail extraction or validation.

When should I use a browser renderer?

Use one when the required content is produced after JavaScript execution or when you need a visual artifact. Keep browser waits bounded and retain diagnostics so rendering failures are distinguishable from network failures.

Final operating checklist

  • Data question, schema, freshness, and validation rules are written down.
  • API and bulk-download options were assessed.
  • robots.txt, terms, and authorization requirements were reviewed.
  • User-agent identifies the project and contact route.
  • URL discovery, canonicalization, and scope limits are explicit.
  • Each host has a conservative limiter with jitter.
  • 429 and 5xx responses trigger backoff; persistent 403s trigger investigation and a stop.
  • Conditional requests and caching are used where appropriate.
  • Redirects, terminal statuses, retries, and dead-letter items are recorded.
  • Extraction validates required fields and versions parsers.
  • Metrics cover coverage, errors, latency, bytes, and data quality.
  • Every record carries provenance, retrieval time, hash, and version metadata.

Reliable crawling is an operating discipline: limit what you request, respect the site, react to its health, validate what you collect, and preserve enough evidence to reproduce the result.