How to Build AI-Ready Web Crawlers in Python
Build a permission-aware Scrapy crawler that produces clean, provenance-rich records for search, RAG, and LLM workflows.

Direct answer: Build an AI-ready crawler as a staged, permission-aware pipeline. Use Scrapy for scheduling, robots.txt checks, duplicate filtering, retries, and extraction; define a document schema before writing selectors; normalize content and metadata; validate every page type; and keep provenance so an embedding, search result, or model answer can be traced back to its source. Add Playwright only when the required content is missing from the HTTP response.
This guide shows a complete Python implementation pattern, including access policy, canonical URLs, clean extraction, JavaScript escalation, validation, storage, observability, and RAG preparation.
1. Define the crawl contract before writing code
A crawler becomes difficult to operate when its rules are hidden in callbacks. Write a crawl contract first. It should answer:
- Which domains and URL prefixes are allowed?
- Which paths, query parameters, file types, and languages are excluded?
- What are the maximum depth, concurrency, delay, timeout, and retry limits?
- How are canonical URLs, redirects, fragments, and duplicate documents handled?
- Which fields are required for indexing?
- How long are raw responses and normalized records retained?
Use a record shape that treats every page as a document with provenance:
{
"url": "https://example.com/page",
"canonical_url": "https://example.com/page",
"retrieved_at": "2026-09-29T08:46:25Z",
"published_at": "2026-09-01",
"title": "Page title",
"author": "Author name",
"site_name": "Example",
"language": "en",
"content_markdown": "# Clean page content",
"headings": ["Introduction"],
"links": [],
"status": 200,
"content_type": "text/html",
"content_hash": "sha256:...",
"parser_version": "site-parser-1",
"extraction_status": "ok",
"extraction_warnings": []
}
Keep the original URL, canonical URL, retrieval time, parser version, and content hash on every record. Those fields make deduplication, citation, incremental refreshes, and parser rollbacks possible.
2. Make robots.txt and access policy a hard gate
Fetch and evaluate robots.txt before scheduling requests. Give the crawler a descriptive user-agent with a contact URL or email, obey disallow rules and crawl delays where supplied, and stop on access challenges. Scrapy documents spiders as classes that control link following and structured item extraction, and its overview covers selectors, feed exports, robots.txt support, and storage backends (spider documentation, Scrapy overview).

OpenAI distinguishes OAI-SearchBot, used to surface sites in ChatGPT search, from GPTBot, associated with training use; publishers can control them independently in robots.txt (OpenAI crawler documentation). Robots changes can take about 24 hours to affect search systems. A robots.txt allowance does not bypass authentication, WAF rules, CAPTCHAs, geo restrictions, or terms of service. OpenAI’s guidance explains that these controls can block otherwise legitimate crawlers (crawler access guidance).
Configure Scrapy’s robots middleware and identify your crawler:
# settings.py
ROBOTSTXT_OBEY = True
ROBOTSTXT_USER_AGENT = 'AcmeResearchBot/1.0 (+https://example.com/bot)'
USER_AGENT = 'AcmeResearchBot/1.0 (+https://example.com/bot)'
CONCURRENT_REQUESTS = 8
DOWNLOAD_DELAY = 0.5
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 0.5
AUTOTHROTTLE_MAX_DELAY = 10
RETRY_HTTP_CODES = [408, 425, 429, 500, 502, 503, 504]
DOWNLOAD_TIMEOUT = 30
Handle 401, 403, 429, and challenge pages as explicit outcomes. Do not rotate identities or brute-force retries to evade a site’s controls.
3. Create a Scrapy project with typed items
Install the crawler and an extraction library:
python -m venv .venv
source .venv/bin/activate
pip install scrapy trafilatura w3lib
scrapy startproject ai_crawler
cd ai_crawler
Define an item so missing fields are visible during validation:
# ai_crawler/items.py
import scrapy
class PageItem(scrapy.Item):
url = scrapy.Field()
canonical_url = scrapy.Field()
retrieved_at = scrapy.Field()
published_at = scrapy.Field()
updated_at = scrapy.Field()
title = scrapy.Field()
author = scrapy.Field()
site_name = scrapy.Field()
language = scrapy.Field()
content_markdown = scrapy.Field()
headings = scrapy.Field()
links = scrapy.Field()
status = scrapy.Field()
content_type = scrapy.Field()
content_hash = scrapy.Field()
parser_version = scrapy.Field()
extraction_status = scrapy.Field()
extraction_warnings = scrapy.Field()
Keep discovery, fetching, extraction, and indexing separable. A failed embedding request should not require downloading the page again.
4. Build a permission-aware spider
The spider below follows only approved hosts, removes fragments before scheduling, limits depth, and emits normalized records. Scrapy’s duplicate filter handles repeated requests, while the explicit canonicalization function prevents common query-string duplicates.
# ai_crawler/spiders/site.py
from datetime import datetime, timezone
from hashlib import sha256
from urllib.parse import urlsplit, urlunsplit
import scrapy
import trafilatura
from w3lib.url import canonicalize_url
from ai_crawler.items import PageItem
class SiteSpider(scrapy.Spider):
name = 'site'
allowed_domains = ['example.com']
start_urls = ['https://example.com/']
custom_settings = {'DEPTH_LIMIT': 3}
def normalize_url(self, url):
parts = urlsplit(canonicalize_url(url, keep_fragments=False))
return urlunsplit((parts.scheme, parts.netloc.lower(), parts.path or '/', parts.query, ''))
def parse(self, response):
if response.status != 200:
self.logger.warning('Skipping %s with status %s', response.url, response.status)
return
content_type = response.headers.get('Content-Type', b'').decode('latin1').lower()
if 'text/html' not in content_type:
return
html = response.text
extracted = trafilatura.extract(
html, output_format='markdown', include_links=True,
include_tables=True, include_comments=False, include_images=False
)
metadata = trafilatura.extract_metadata(html)
text = extracted or ''
canonical = response.css("link[rel='canonical']::attr(href)").get()
canonical_url = response.urljoin(canonical) if canonical else self.normalize_url(response.url)
title = metadata.title if metadata else response.css('title::text').get()
warnings = [] if text.strip() else ['empty_extraction']
yield PageItem(
url=response.url,
canonical_url=self.normalize_url(canonical_url),
retrieved_at=datetime.now(timezone.utc).isoformat(),
published_at=getattr(metadata, 'date', None) if metadata else None,
updated_at=None,
title=title.strip() if title else None,
author=getattr(metadata, 'author', None) if metadata else None,
site_name=getattr(metadata, 'sitename', None) if metadata else None,
language=response.css('html::attr(lang)').get(),
content_markdown=text,
headings=response.css('h1, h2, h3::text').getall(),
links=[response.urljoin(h) for h in response.css('a::attr(href)').getall()],
status=response.status,
content_type=content_type,
content_hash='sha256:' + sha256(text.encode('utf-8')).hexdigest(),
parser_version='site-parser-1',
extraction_status='ok' if text.strip() else 'quarantined',
extraction_warnings=warnings,
)
for href in response.css('a::attr(href)').getall():
next_url = self.normalize_url(response.urljoin(href))
if urlsplit(next_url).netloc in self.allowed_domains:
yield scrapy.Request(next_url, callback=self.parse)
Run it as JSON Lines for a replayable intermediate artifact:
scrapy crawl site -O pages.jl
For production, replace the example domain, add path allowlists, reject downloads such as archives and videos, and store redirect chains and response headers where they are needed for audits.
5. Extract content that works for AI systems
Raw HTML contains navigation, ads, cookie notices, repeated headers, and scripts. Trafilatura can produce Markdown and metadata such as title, author, date, and site name (Scrapy extraction guide). Preserve headings, lists, tables, code blocks, captions, and link targets when they carry meaning. Article-focused extraction may return little or nothing for product pages, category pages, and listings, so use page-type-specific parsers and retain the raw HTML or a content hash when reproducibility matters.
Clean and normalize before chunking. Attach document metadata to every chunk:
chunk = {
'text': section_text,
'url': record['url'],
'canonical_url': record['canonical_url'],
'title': record['title'],
'retrieved_at': record['retrieved_at'],
'published_at': record['published_at'],
'content_hash': record['content_hash'],
'parser_version': record['parser_version']
}
Do not send records with empty bodies, missing canonical URLs, implausible lengths, or parser warnings directly to embeddings. Put them in a quarantine stream for review.
6. Escalate to a browser only for JavaScript-dependent pages
First inspect the HTTP response. Scrapy’s dynamic-content documentation notes that useful data may be embedded in JavaScript or loaded from an external resource, and recommends checking the response obtained by an HTTP client before assuming a browser is required (dynamic-content documentation).

Use scrapy-playwright when the content appears only after JavaScript execution, scrolling, interaction, or client-side requests. Keep browser requests narrow because they add CPU, memory, timing, and failure modes. Prefer an accessible JSON endpoint or embedded state object when one exists and its use is permitted.
pip install scrapy-playwright playwright
playwright install chromium
# settings.py additions
DOWNLOAD_HANDLERS = {
'http': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
'https': 'scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler',
}
TWISTED_REACTOR = 'twisted.internet.asyncioreactor.AsyncioSelectorReactor'
yield scrapy.Request(
url,
meta={
'playwright': True,
'playwright_page_methods': [
{'method': 'wait_for_selector', 'args': ['main']},
],
},
callback=self.parse,
)
Set a strict browser timeout, close pages promptly, and record whether a record came from HTTP or browser rendering. A browser should be an escalation policy, not the default path for every URL.
7. Validate page variants before indexing
Create fixtures for every important template: article, documentation page, product page, listing, login wall, error page, and JavaScript-rendered page. Validate required fields, title and date parsing, canonical URLs, body length, link extraction, boilerplate removal, and content type. Scrapy’s AI workflow recommends defining a schema, downloading representative pages, comparing variants, validating the extraction specification, and generating runnable tests (Scrapy AI workflow).
Add drift alarms for sudden changes in status codes, empty-body rates, null fields, duplicate ratios, and content-length distributions. Quarantine failures instead of indexing them. Store parser version and crawl timestamp so a corrected parser can rebuild the index.
def validate(record):
errors = []
if not record.get('canonical_url'): errors.append('missing_canonical')
if not record.get('title'): errors.append('missing_title')
if len((record.get('content_markdown') or '').strip()) < 200:
errors.append('content_too_short')
if record.get('extraction_status') != 'ok':
errors.append('extraction_not_ok')
return errors
8. Handle failures, freshness, and scale
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or challenge HTML | WAF, bot mitigation, authentication, or disallowed access | Stop, verify permission and terms, slow the crawl, and request an approved access method. |
| 429 responses | Too much concurrency or a published quota | Reduce concurrency, honor Retry-After, increase delay, and schedule incremental crawls. |
| Empty extraction | Wrong page type, boilerplate-heavy layout, or client-rendered content | Inspect the response, add a page-type parser, or escalate that URL to a browser. |
| Duplicate documents | Tracking parameters, fragments, redirects, or alternate canonicals | Canonicalize URLs, honor rel=canonical, and deduplicate by canonical URL plus content hash. |
| Stale RAG answers | Old records or chunks without timestamps | Use incremental recrawls, compare hashes, and filter or boost by freshness. |
| Parser drift | Template or markup change | Use fixtures, null-rate alarms, parser versions, and a quarantine queue. |
For performance, keep HTTP concurrency modest, enable AutoThrottle, avoid rendering pages that expose the needed data in HTML, and cache during development. For reliability, separate discovery, fetching, extraction, validation, and indexing so each stage can retry independently. Operating cost comes from network volume, browser CPU, proxy use, storage, and managed services. Measure those dimensions before scaling.
9. Or skip the browser setup
If your crawler only needs a clean visual capture of a page, ScreenshotNeo provides a single GET endpoint that returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
See the ScreenshotNeo API documentation for all options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options cover full-page captures with lazy images, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS rendering, custom JavaScript and CSS, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs, and a usage API. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is included on every plan. Create a free ScreenshotNeo account.
10. Prepare records for search and RAG
- Normalize whitespace and character encoding.
- Preserve semantic structure and links.
- Split by headings or logical sections before applying a token limit.
- Attach URL, title, canonical URL, dates, hash, and parser version to every chunk.
- Embed only records that pass validation.
- Store source references so answers can cite the page and retrieval time.
- Re-crawl incrementally using canonical URLs and content hashes.
When a parser changes, reprocess quarantined and affected records before deleting the old index. Keep crawl runs identifiable so you can compare extraction quality over time.
Frequently asked questions
Should I start with Scrapy or Playwright?
Start with Scrapy. Add Playwright only for content that is absent from the HTTP response or requires interaction.
Can robots.txt permission guarantee access?
No. Authentication, WAFs, CAPTCHAs, rate limits, geo rules, and terms can still restrict access.
How do I prevent bad pages from entering my vector index?
Validate required fields and body length, monitor drift metrics, and quarantine records with extraction warnings before embedding.
What metadata is essential for citations?
At minimum, keep the source URL, canonical URL, retrieval timestamp, title, content hash, and parser version.
When should I use a hosted browser or crawler service?
Consider one when JavaScript coverage, proxy rotation, monitoring, or managed deployment exceeds what your local Scrapy workers can reliably operate. Verify current terms and compliance requirements before adoption.


