ScreenshotNeo

BlogEngineering

Enterprise Web Crawler FAQs

Design an enterprise web crawler that respects robots.txt, scales safely, handles JavaScript, deduplicates URLs, and feeds a reliable search index.

By the ScreenshotNeo team1 October 202612 min read

An enterprise web crawler is an operational pipeline, not a script that downloads HTML. URLs enter a queue, scope and robots rules are checked, host-specific fetches are scheduled, responses are retried or discarded according to policy, content is extracted and deduplicated, and accepted documents are sent to a search index or knowledge base.

The direct design answer is:

  1. Define ownership, authorization, allowed hosts, data handling, and retention before crawling.
  2. Discover URLs from seed lists, sitemaps, links, and approved feeds.
  3. Fetch and cache each host’s /robots.txt; treat it as a cooperation protocol, never as authentication.
  4. Use a per-host scheduler with concurrency limits, backoff, status-code handling, and a clear crawler user-agent.
  5. Normalize and deduplicate URLs before and after fetching.
  6. Render JavaScript only when required, and record the rendering mode used.
  7. Extract canonical content, metadata, links, and change fingerprints.
  8. Send versioned documents to the index, and process deletions as well as additions and changes.
  9. Measure queue age, host rate, response classes, retries, duplicates, ingestion success, freshness, and robots compliance.

RFC 9309 defines the Robots Exclusion Protocol. It explicitly says: "These rules are not a form of access authorization." Private pages still need authentication or another application-layer control.

1. What is an enterprise web crawler?

An enterprise crawler continuously discovers and retrieves web resources for a defined business purpose, such as internal search, a knowledge base, documentation indexing, or content monitoring. It normally contains these services:

Component Responsibility
Seed manager Accepts approved domains, URL lists, feeds, and sitemap locations.
URL frontier Stores pending, leased, completed, failed, and delayed URLs.
Policy service Checks scope, robots rules, authentication policy, exclusions, and data classifications.
Host scheduler Applies per-host concurrency, delay, quotas, and adaptive backoff.
Fetcher Makes HTTP requests, follows permitted redirects, records headers and status, and enforces size and timeout limits.
Renderer Runs a browser for pages whose useful content or links require JavaScript.
Parser and extractor Extracts main content, title, metadata, links, structured data, and page-level robots directives.
Deduplication service Normalizes URLs and compares content fingerprints to avoid duplicate work.
Ingestion sink Writes documents, embeddings, or files to the destination index or knowledge base.
Observability Reports queue depth, latency, errors, retries, freshness, and policy decisions.

2. Reference crawl pipeline

A useful mental model is:

seeds and sitemaps
        |
        v
URL normalization and scope checks
        |
        v
robots.txt policy cache
        |
        v
per-host queue and rate limiter
        |
        v
HTTP fetch ----> optional browser rendering
        |
        v
status, headers, robots meta, canonical URL
        |
        v
content extraction and link discovery
        |
        v
fingerprint, deduplicate, version
        |
        v
search index or knowledge base

Keep policy decisions close to the fetch boundary. A URL can be discovered legally but still be outside the configured scope, disallowed by robots rules, unauthenticated, too large, or subject to a retention restriction.

Use a durable queue with leases. A worker should claim a URL for a limited period; if it crashes, the lease expires and another worker can retry it. Store an attempt record separately from the URL record so operational history is not lost when a page changes.

3. Does robots.txt protect private pages?

No. Robots.txt controls crawler access requests; it does not secure a page and does not reliably keep a URL out of search results. The RFC describes robots.txt as a cooperation mechanism and recommends valid application-layer security for restricted content. Google likewise documents that a blocked URL can still appear in search results if it is linked elsewhere; use authentication for private resources and an indexing control such as noindex when exclusion from Google results is the goal. See Google Search Central’s robots.txt documentation.

Implement robots handling as follows:

  1. Request https://host/robots.txt before fetching in-scope URLs.
  2. Parse rules for your exact user-agent and the wildcard group.
  3. Apply the most specific matching allow or disallow rule according to the parser you use.
  4. Honor sitemap locations when you use them for discovery.
  5. Cache a successfully fetched file, but do not use a cached version for more than 24 hours unless the file is unreachable, as recommended by RFC 9309.
  6. Record whether a URL was allowed, disallowed, or deferred because the policy file was unavailable.

RFC 9309 requires a robots parser limit of at least 500 KiB. That is a parsing minimum for robots.txt, not a page-size allowance for crawled content.

4. Minimal crawler implementation in Python

The following example is intentionally small, but it demonstrates the core controls: scope, robots checks, a clear user-agent, a per-request delay, URL normalization, extraction, and a bounded queue. Production systems should replace the in-memory queue with durable storage and add authentication, leases, metrics, and a real HTML parser.

import time
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup

START = 'https://example.com/'
ALLOWED_HOST = urlparse(START).netloc
USER_AGENT = 'AcmeEnterpriseCrawler/1.0 (+https://example.com/crawler-info)'
DELAY_SECONDS = 10
MAX_PAGES = 100

session = requests.Session()
session.headers.update({'User-Agent': USER_AGENT})
robots = RobotFileParser(urljoin(START, '/robots.txt'))
try:
    robots.read()
except Exception:
    # Decide this policy explicitly for your organization. A conservative
    # crawler pauses the host until robots.txt can be retrieved.
    raise RuntimeError('robots.txt could not be fetched')

queue = deque([START])
seen = set()

def normalize(url):
    url, _fragment = urldefrag(url)
    parsed = urlparse(url)
    if parsed.scheme not in ('http', 'https'):
        return None
    if parsed.netloc != ALLOWED_HOST:
        return None
    return url

while queue and len(seen) < MAX_PAGES:
    url = normalize(queue.popleft())
    if not url or url in seen or not robots.can_fetch(USER_AGENT, url):
        continue
    seen.add(url)
    try:
        response = session.get(url, timeout=(10, 30), allow_redirects=True)
    except requests.RequestException as exc:
        print('fetch error', url, exc)
        continue
    if response.status_code == 429:
        time.sleep(60)
        continue
    if response.status_code != 200:
        print('skip', response.status_code, url)
        continue
    content_type = response.headers.get('content-type', '')
    if 'text/html' not in content_type:
        continue
    soup = BeautifulSoup(response.text, 'html.parser')
    title = soup.title.get_text(' ', strip=True) if soup.title else ''
    text = soup.get_text(' ', strip=True)
    print({'url': response.url, 'title': title, 'characters': len(text)})
    for link in soup.select('a[href]'):
        child = normalize(urljoin(response.url, link['href']))
        if child and child not in seen:
            queue.append(child)
    time.sleep(DELAY_SECONDS)

This sample uses one request every 10 seconds. AWS Prescriptive Guidance gives one request every 10–15 seconds as an example for small or medium-sized sites, and 1–2 requests per second for larger sites or sites with explicit permission. Those figures are contextual guidance, not universal limits. Use the site’s instructions and your authorization as the controlling inputs.

5. URL discovery, normalization, and deduplication

Start with the smallest complete source of URLs. Prefer approved seed lists and XML sitemaps, including sitemap locations referenced by robots.txt. Discover additional links only inside the configured scope.

Normalize before queue insertion:

  • Remove URL fragments unless the application treats them as distinct resources.
  • Lowercase the host and remove default ports.
  • Resolve relative links against the final response URL.
  • Apply a documented policy for trailing slashes, default documents, and duplicate query parameters.
  • Drop tracking parameters such as campaign tags only when doing so is safe for the target site.
  • Preserve query parameters that change content.
  • Reject unsupported schemes, private network destinations, and hosts outside the allowlist.

Use at least two duplicate checks: a URL key before fetch and a content fingerprint after extraction. Two URLs can return identical content, while one URL can change over time. Store a canonical URL, an HTTP entity tag or last-modified value when provided, a normalized-content hash, and the crawl timestamp.

6. Scheduling, retries, and JavaScript rendering

Schedule by host rather than using one global worker pool. A global limit can still overload a small site if many workers happen to select it simultaneously. Each host queue should track next-allowed time, active requests, recent status codes, and backoff state.

Response Typical action
200 Parse, extract, fingerprint, and enqueue permitted links.
3xx Follow only allowed redirects; re-check scope and robots policy for the destination.
401/407 Do not guess credentials. Check the configured authentication flow and secret expiry.
403 Pause or stop the host when denials continue; investigate authorization and policy.
404 Mark the URL absent; retain history if the destination index must process deletions.
429 Reduce the host rate, honor Retry-After when present, and retry with backoff.
5xx or timeout Retry a bounded number of times with exponential backoff and jitter, then quarantine the URL.

Use a browser only when plain HTTP cannot obtain the required content or links. Browser rendering costs more CPU and memory, has more failure modes, and can execute third-party code. Set a navigation timeout, restrict outbound destinations, block unnecessary resource types, and capture the final HTML after the required interaction. JavaScript-generated links may require simulated interaction; some managed crawlers document that such links are not discovered automatically.

7. Extraction and ingestion into a knowledge base

Extract a stable document envelope rather than sending raw HTML directly to the index:

{
  'id': 'https://example.com/docs/page',
  'canonical_url': 'https://example.com/docs/page',
  'title': 'Page title',
  'text': 'Cleaned main content',
  'language': 'en',
  'published_at': None,
  'modified_at': None,
  'source_status': 200,
  'content_hash': 'sha256:...',
  'crawled_at': '2026-10-01T12:00:00Z',
  'links': ['https://example.com/docs/next'],
  'policy': {'robots_allowed': True}
}

Keep source URL, crawl timestamp, content hash, and deletion state in the index. On an incremental crawl, process added, changed, and deleted content. If a URL disappears from a sitemap, that is a discovery signal, not proof of deletion; confirm with a fetch or an authoritative feed before removing an indexed document.

Chunking and embedding are downstream choices. Preserve headings and source links so retrieval results can be traced back to the page. Store raw or normalized content only for the retention period approved by your organization.

8. Authentication, privacy, and network security

Robots.txt does not grant permission to crawl. Crawl only sites your organization owns or has explicit authorization to access. Define which credentials may be used, where they are stored, how they rotate, and which pages may enter the index.

  • Keep credentials in a secret manager and inject them at fetch time.
  • Do not place cookies, bearer tokens, or passwords in URLs or logs.
  • Use outbound-only network access for crawler workers where possible.
  • Block requests to loopback, link-local, and private-network addresses to reduce server-side request forgery risk.
  • Apply response-size, decompression, redirect-count, and execution-time limits.
  • Classify extracted data and enforce access controls in the destination index.
  • Log policy decisions without logging sensitive page contents by default.

9. Build versus managed crawler

Build when you need unusual authentication, bespoke rendering, specialized extraction, or tight control over network and storage. A managed crawler can reduce queue, retry, scheduling, and incremental-sync maintenance.

Decision axis Questions to answer
Permission and robots Can the system enforce your authorization and robots policy per host?
Authentication Can secrets, sessions, and login changes be handled safely?
Rendering Are JavaScript interactions and client-side links required?
Rate control Can you set host-level limits and respond to 429 or 403 responses?
Refresh Are additions, changes, and deletions synchronized incrementally?
Limits What file sizes, page counts, timeouts, and attachment types are supported?
Integration Does output reach your index or knowledge base without custom transport?
Security Where does content run, and how are credentials and logs isolated?
Operations Who owns alerts, retries, upgrades, and recovery?
Total cost What are service charges, browser compute, storage, and engineering time?

Amazon Bedrock Web Crawler is one documented managed example. AWS describes an initial full sync followed by incremental syncs, retries, URL deduplication, crawler identification, robots directives, and page-level robots meta tags. Its documentation also describes limitations involving interaction-driven JavaScript discovery, authentication failures, 429 responses, and file-size limits. Verify current service behavior and limits before choosing it, and follow AWS requirements to crawl only sites you own or are authorized to crawl.

10. Troubleshooting common crawler failures

Symptom Likely cause Fix
Every URL is skipped Host scope, scheme, or robots matching is wrong. Log the normalized URL, matched rule, user-agent, and scope decision.
Pages appear in search but not the index Discovery depends on JavaScript interaction or an unseeded route. Add sitemap or seed URLs, or use a controlled browser workflow.
Many 429 responses Per-host rate is too high. Honor Retry-After, reduce concurrency, and add jittered backoff.
Persistent 403 responses The crawler is not authorized or is being denied. Stop the host and confirm permission; do not rotate identities to evade controls.
401 after a successful run Expired credentials or changed login configuration. Rotate the secret, validate the session flow, and alert before retrying broadly.
Duplicate documents Query parameters, fragments, redirects, or aliases were not normalized. Canonicalize URLs and compare normalized-content hashes.
Large pages fail Response or attachment size limit. Set explicit limits, skip unsupported files, or use an approved file connector.
Fresh content is missing Refresh interval is longer than the publication cycle. Use change feeds, sitemaps, conditional requests, or a shorter targeted recrawl.
Worker retries forever No retry budget or dead-letter state. Cap attempts, record the reason, and quarantine the URL for review.

11. Performance and reliability

Scale the queue and workers independently from the host scheduler. More workers do not make one host safe to crawl faster. Measure:

  • queue depth and oldest queued URL;
  • requests per host and active leases;
  • latency by status class and rendering mode;
  • 429, 403, timeout, and 5xx rates;
  • retry volume and dead-letter count;
  • duplicate URL and duplicate-content rates;
  • fetched pages versus successfully ingested pages;
  • document age and freshness by source;
  • robots-rule compliance and policy-decision counts.

Reliability comes from idempotent writes, durable leases, bounded retries, deterministic URL normalization, and replayable extraction. Separate transient failures from permanent policy decisions so a temporary outage does not erase a valid URL and a denial is not retried indefinitely.

12. Cost planning

Estimate total cost as fetch compute plus browser compute plus storage, queueing, indexing, embeddings, bandwidth, observability, and engineering operations. Browser rendering is usually the most resource-intensive path, so classify URLs and render only those that need it. Incremental crawls reduce work when you have trustworthy change signals.

Do not use request-rate examples as capacity guarantees. Measure your own pages, authorization constraints, response sizes, rendering requirements, and destination indexing cost. A managed service may reduce maintenance while adding usage charges and product-specific limits; a custom crawler may reduce per-request fees while increasing engineering and operational burden.

13. Or skip the browser setup

For individual page screenshots used in documentation, audits, visual checks, or agent workflows, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Features include full-page capture with lazy images loaded, CSS element capture, dark mode, device presets and custom viewports, retina scale, PDF paper sizes and page ranges, HTML or CSS rendering, custom JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, usage reporting, and an OpenAPI specification. Plans include 1,000 free shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

14. Enterprise crawler FAQ

How often should robots.txt be refreshed?

RFC 9309 recommends not using a cached file for more than 24 hours unless it is unreachable. A shorter refresh can be appropriate for high-risk or frequently changing hosts.

Can a crawler ignore robots.txt with permission?

Permission and robots policy are separate decisions. Document the authorization and configure the crawler’s policy explicitly; do not assume that permission removes security, privacy, or contractual obligations.

Should every page be rendered in a browser?

No. Fetch static HTML first and render only when required content or links are produced by JavaScript or interaction.

What is the safest retry policy?

Use bounded retries, exponential backoff with jitter, host-level limits, and a dead-letter state. Pause on 429 and investigate continuing 403 responses.

How do I prevent stale results?

Track crawl timestamps and content hashes, use conditional requests or change feeds where available, run targeted recrawls, and process deletions explicitly.

When should I choose a managed crawler?

Choose one when its authentication, rendering, limits, integrations, security model, and refresh behavior match your workload and the maintenance savings justify its cost. Verify current documentation before committing.