Web Crawlers Explained: How to Crawl a Website
Learn how crawlers discover and fetch pages, how robots.txt and sitemaps guide them, and how to build a small, responsible crawler.
A web crawler is software that discovers URLs and fetches the resources at those addresses. To crawl a small site, start with seed URLs, fetch each eligible URL once, extract links, normalize and deduplicate them, enforce scope and request limits, and add eligible discoveries to a queue. Crawling is only the fetch stage: a search engine may crawl a page without indexing it or showing it in results.
1. What a web crawler does
A crawler is also called a bot, robot, or spider. It automatically requests web resources and processes their responses for a purpose such as discovering pages, checking links, or collecting content. There is no central registry of every page on the web. Search engines build lists of known URLs from pages they already know, links they find, and submitted sitemaps, then choose which URLs to fetch. Google’s guide to how Search works describes discovery, crawling, indexing, and serving as distinct stages.
2. The crawl loop
A useful mental model for a small crawler is a queue of URLs waiting to be fetched and a set of URLs already seen. The exact architecture varies; this loop is an implementation model, not a universal crawler design.
- Choose seed URLs. These may be a homepage, known page URLs, or URLs from a sitemap.
- Queue eligible URLs. Normalize URLs, reject out-of-scope addresses, and avoid adding duplicates.
- Fetch one URL. Send an HTTP request with an appropriate user agent. Handle status codes, timeouts, redirects, and response size limits.
- Process the response. Check its content type and parse HTML when appropriate.
- Discover links. Resolve relative links against the page URL, normalize them, and apply scope and access rules.
- Limit the work. Use a page cap, concurrency limit, delay or backoff, and a clear stopping condition.
For search crawlers, fetching is not the same as indexing. Indexing systems decide whether and how crawled content is stored, and serving systems decide what appears for a query. A successfully fetched URL is not guaranteed to be indexed. Google Search Central explains these separate stages.
3. Build a small crawler in Python
This runnable example crawls HTML pages on one host. It honors robots.txt disallow rules for the declared user agent, limits page count and request rate, follows ordinary links, and avoids revisiting normalized URLs. It is intentionally a small single-process crawler, not a production search engine.
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.robotparser import RobotFileParser
import time
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/"
USER_AGENT = "ExampleLearningCrawler/1.0 (contact: crawler@example.com)"
MAX_PAGES = 50
DELAY_SECONDS = 1.0
TIMEOUT_SECONDS = 15
MAX_RESPONSE_BYTES = 2_000_000
start = urldefrag(START_URL)[0]
origin = urlparse(start)
robots_url = f"{origin.scheme}://{origin.netloc}/robots.txt"
robots = RobotFileParser()
robots.set_url(robots_url)
try:
robots.read()
except Exception:
# A real crawler should define an explicit policy for robots.txt failures.
# This example stops rather than assuming permission.
raise RuntimeError(f"Could not retrieve {robots_url}; define your robots policy")
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})
queue = deque([start])
seen = {start}
last_request_at = 0.0
while queue and len(seen) <= MAX_PAGES:
url = queue.popleft()
if not robots.can_fetch(USER_AGENT, url):
print("ROBOTS_BLOCKED", url)
continue
wait = DELAY_SECONDS - (time.monotonic() - last_request_at)
if wait > 0:
time.sleep(wait)
try:
response = session.get(url, timeout=TIMEOUT_SECONDS, allow_redirects=True, stream=True)
last_request_at = time.monotonic()
response.raise_for_status()
content_type = response.headers.get("Content-Type", "").lower()
if "text/html" not in content_type:
print("SKIP_NON_HTML", response.status_code, url, content_type)
continue
body = bytearray()
for chunk in response.iter_content(chunk_size=16_384):
body.extend(chunk)
if len(body) > MAX_RESPONSE_BYTES:
raise ValueError("response exceeded size limit")
final_url = urldefrag(response.url)[0]
soup = BeautifulSoup(body, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else "(no title)"
print("PAGE", response.status_code, final_url, title)
for anchor in soup.select("a[href]"):
candidate = urldefrag(urljoin(final_url, anchor["href"]))[0]
parsed = urlparse(candidate)
if parsed.scheme not in ("http", "https"):
continue
if parsed.netloc.lower() != origin.netloc.lower():
continue
if candidate not in seen and robots.can_fetch(USER_AGENT, candidate):
seen.add(candidate)
queue.append(candidate)
except requests.RequestException as exc:
print("FETCH_ERROR", url, str(exc))
except ValueError as exc:
print("RESPONSE_LIMIT", url, str(exc))
print("Discovered URLs:", len(seen))
Install the two dependencies with python -m pip install requests beautifulsoup4, save the code as crawler.py, and run python crawler.py. Replace the example domain and contact value before using it. The robots.txt handling here stops on retrieval errors as a conservative example policy; production behavior should be chosen deliberately and documented.
Why these safeguards matter
- Scope: This example stays on the seed host. A subdomain is a different host under this check; broaden scope only when intended.
- Deduplication: Removing URL fragments avoids crawling the same HTTP resource under different in-page anchors. Query strings remain distinct because they can represent distinct pages, but they can also create URL explosions.
- Rate control: A one-second delay is an example setting, not a universal safe rate. Reduce concurrency and back off on server errors or rate limits.
- Resource limits: Timeouts and response byte caps prevent a crawl from waiting forever or consuming unbounded memory.
- Redirects: The final response URL is used as the base for link resolution. For a production crawler, also cap redirect hops and record redirect destinations.
4. Crawl a page with cURL, Python, or Node.js
For a one-off fetch, these examples retrieve a page but do not implement a crawler queue, link discovery, scope rules, or robots policy. Use them to inspect a response, then add the crawl loop and safeguards when fetching multiple URLs.
cURL
curl --fail --location --max-time 15 \
-A "ExampleLearningCrawler/1.0 (contact: crawler@example.com)" \
-H "Accept: text/html" \
https://example.com/ -o page.html
Python
import requests
response = requests.get(
"https://example.com/",
headers={"User-Agent": "ExampleLearningCrawler/1.0 (contact: crawler@example.com)"},
timeout=15,
)
response.raise_for_status()
print(response.status_code, response.url, response.headers.get("Content-Type"))
print(response.text[:500])
Node.js
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 15_000);
try {
const response = await fetch('https://example.com/', {
headers: {
'User-Agent': 'ExampleLearningCrawler/1.0 (contact: crawler@example.com)',
'Accept': 'text/html',
},
signal: controller.signal,
redirect: 'follow',
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(response.status, response.url, response.headers.get('content-type'));
console.log((await response.text()).slice(0, 500));
} finally {
clearTimeout(timer);
}
These are ordinary HTTP fetches. They do not execute page JavaScript. A plain fetch is often enough for static HTML; browser rendering is needed only when the information your task depends on is added after scripts run.
5. Robots.txt: what it controls and what it does not
The Robots Exclusion Protocol (REP), usually published at /robots.txt, lets site owners express which URL paths compliant crawlers may fetch. Google says it fetches and parses the file before crawling. The file must be at the top level and applies to the matching scheme, host, and port. See Google’s robots.txt specification notes and the REP standard, RFC 9309.
User-agent: *
Disallow: /private-preview/
Allow: /private-preview/public-sample.html
Sitemap: https://example.com/sitemap.xml
Supported Google directives include User-agent, Allow, Disallow, and Sitemap. Google does not support Crawl-delay; do not rely on that field to manage Googlebot pacing. Rules are interpreted by user-agent group and path matching, so check the target crawler’s documentation and test the exact paths you intend to control.
| Goal | Use | Important limit |
|---|---|---|
| Ask compliant crawlers not to fetch a path | robots.txt disallow rule | Does not provide privacy or authentication |
| Keep private content private | Authentication or access control | A robots rule is not a security boundary |
| Prevent eligible content from appearing in Google Search | Use an appropriate mechanism such as noindex or password protection |
A crawler must be able to fetch a page to see a page-level noindex directive |
A disallowed URL can still appear in search results if other pages link to it, even when its content was not fetched. For exclusion guidance, see Google’s robots.txt guide. Never put sensitive data behind a robots rule.
6. Discover URLs with links and sitemaps
Links are a natural discovery path: a crawler fetches a known page, extracts its links, and considers those URLs. Resolve relative links against the final page URL, accept only schemes your crawler supports, and then enforce host, path, robots, and duplicate rules.
An XML sitemap is another way to suggest URLs for discovery. It is a list for crawlers to consider, not a guarantee that every URL will be fetched or indexed. Keep it current and include useful lastmod values for pages that changed. The Sitemaps Protocol describes the file format; Google’s crawl budget guidance recommends maintaining sitemaps and accurate modification dates.
For a custom crawler, decide whether sitemap URLs expand your crawl scope. A sitemap can contain URLs outside the path you initially expected, and malformed or hostile input should not be trusted blindly.
7. Crawl budget, capacity, and URL traps
Google describes crawl budget as the URLs it can and wants to crawl. Crawl capacity concerns how much fetching a host can tolerate; crawl demand concerns which URLs are worth revisiting and can vary with site size, update frequency, quality, relevance, popularity, URL inventory, and staleness. There is no single crawl rate or threshold appropriate for every website. See Google’s crawl budget documentation.
For your own crawler, capacity is a practical resource limit: how quickly the site can respond without harm and how much bandwidth, memory, and processing you can spend. Use a small concurrency limit, respect rate-limit responses, apply exponential backoff with jitter to transient failures, and stop or slow down when the server shows distress.
Common URL patterns can create huge or effectively infinite crawl spaces:
- Faceted navigation with arbitrary filter combinations
- Sorting and pagination parameters that generate many equivalent views
- Unrestricted calendars and date ranges
- Session IDs and tracking parameters
- Duplicate URLs caused by case, trailing slash, or query ordering differences
- Malformed relative links that expand into unexpected paths
- Long redirect chains and redirect loops
Google recommends consolidating duplicate pages, limiting redundant URL variants, keeping sitemaps fresh, avoiding long redirect chains, and returning 404 or 410 for permanently removed pages. See its URL structure guidance.
8. HTML fetching versus JavaScript rendering
A basic crawler downloads HTML and parses it without running scripts. It can miss text and links that appear only after client-side JavaScript executes. Google says its crawler renders pages and executes JavaScript; whether your own crawler needs a browser depends on the target pages and your task. Browser rendering adds time, compute cost, and operational complexity, so use it when the raw response lacks the content you need. Google’s JavaScript SEO basics explains its rendering approach.
To determine whether rendering is needed, compare the HTML from an HTTP fetch with what a browser displays. If important links or content are absent from the response, render selectively. Keep URL discovery and page fetching separate so a rendered page does not silently expand the crawl beyond the intended scope.
9. Site-owner checklist for discoverability
- Publish important pages through ordinary crawlable links.
- Maintain a sitemap containing canonical, useful URLs and accurate modification dates.
- Check that robots.txt is at the site root and its rules match the intended host and protocol.
- Use authentication for private information; robots.txt is public and advisory.
- Return correct status codes, including 404 or 410 for permanently removed content.
- Reduce duplicate URL variants, unnecessary parameters, and redirect chains.
- Inspect server logs and Search Console reports to understand crawling behavior rather than assuming every discovered URL is fetched.
10. Troubleshooting common crawler problems
| Symptom | Likely cause | Fix |
|---|---|---|
| No links are discovered | The page has no anchors in returned HTML, or links are injected by JavaScript | Inspect the raw response; render only if the needed links require JavaScript |
| The crawler revisits pages endlessly | URL normalization or deduplication is missing; query parameters create variants | Canonicalize carefully, remove known tracking parameters, and define allowed query keys |
| Requests receive 403 or 429 | The server denies access or rate-limits the client | Check permissions and crawler identification; reduce request rate and concurrency, and honor retry guidance |
| Many 5xx responses or slow requests | Server overload, transient failure, or too much concurrency | Back off, lower concurrency, set timeouts, and retry only transient errors with a bounded policy |
| robots.txt appears ignored | Wrong scheme, host, port, path, user-agent group, or a crawler that does not honor REP | Test the exact URL against the exact bot group and verify the file is at that origin’s root |
| A blocked URL appears in search | Its URL is known through links while content fetch is disallowed | Use noindex where fetch is allowed or protect it with authentication; do not treat robots.txt as removal or privacy control |
| Large sites are only partly crawled | Excessive duplicate URLs, weak discovery paths, server limits, or crawl demand priorities | Improve links and sitemap quality, reduce URL traps, remove redirect chains, and inspect server capacity |
| Crawler consumes too much memory | Unbounded queue, response bodies, or saved page content | Set page and byte limits, stream responses, persist the queue, and discard unneeded bodies |
| Internal links are misclassified as external | Relative URL resolution or host comparison is incorrect | Resolve against the final response URL and compare parsed, normalized hostnames |
11. Performance, reliability, and cost
For a small crawl, the largest practical costs are network traffic, response parsing, storage, and any browser rendering. A bounded queue and response-size cap keep a crawl predictable. If you need more throughput, increase concurrency gradually while watching latency, errors, and server feedback; no fixed request rate is safe for every target.
Reliability comes from recording each URL’s outcome, distinguishing permanent failures from transient ones, using bounded retries, and saving crawl state so a process can resume. Avoid retrying 404 responses as though they were temporary. For production jobs, record fetch time, final URL, status, content type, bytes, and error category. Respect site policies and applicable access controls.
For search engines, capacity and demand determine how crawl effort is allocated; for a custom crawler, you set your own page, time, bandwidth, and storage budgets. Browser rendering should be a deliberate cost because it uses more resources than an HTTP fetch.
12. Or skip the browser setup
If your immediate goal is a clean screenshot of a page rather than discovering and crawling a site, ScreenshotNeo provides a website screenshot API and MCP server. It is not a general-purpose crawler: it captures a URL you provide. Its API can return PNG, JPEG, WebP, or PDF, with options such as full-page capture, element capture, device presets, custom CSS and JavaScript, and wait conditions. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/ -o shot.webp
Equivalent Python and Node.js requests are available when you want to integrate capture into an application:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await Bun.write('shot.webp', res);
- Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with verdict and billing information in response headers.
- An MCP server lets AI agents use
take_screenshot,get_page_info, andcapture_pdf. - 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
13. FAQ
Does every crawler follow robots.txt?
No. REP is for compliant crawlers. A custom crawler should choose and document its behavior, and site owners should use access controls for restricted information.
Can crawling alone make a page appear in Google?
No. Crawling is a fetch stage. Indexing and serving are separate decisions.
Should a crawler start from a sitemap or homepage?
Use the entry point that matches the task. A homepage reveals linked navigation; a sitemap can expose a known inventory. Combining them can improve discovery, subject to scope and deduplication.
Do all websites require a headless browser to crawl?
No. Use plain HTTP and HTML parsing when the response contains the information you need. Render pages only when required content is added by JavaScript.
References
- Google Search Central: How Google Search works
- Google: How Google interprets the robots.txt specification
- IETF RFC 9309: Robots Exclusion Protocol
- Google: Robots.txt introduction and guide
- Google: Crawl budget management
- Sitemaps Protocol
- Google: URL structure best practices
- Google: JavaScript SEO basics


