How to Build AI Models for Web Scraping
Build a reliable web extraction pipeline by combining Scrapy, browser rendering where needed, trustworthy labels, and careful model evaluation.
To build an AI model for web scraping, define the fields you need, collect permitted pages with a crawler, label examples against a fixed schema, train an extractor or classifier, and evaluate it on pages and domains it has not seen. In most projects, Scrapy should handle crawling and data pipelines; the model should handle ambiguous page classification, extraction, deduplication, or normalization. Use a browser such as Playwright only when the required content is genuinely unavailable in the underlying HTTP response or request.
This guide builds that system as a staged, auditable pipeline. It does not train a general-purpose language model from scratch: that is rarely necessary for a site-specific extraction task.
1. Define the task before crawling
Write down the output schema and what counts as a correct result before gathering data. For a product-page extractor, that could mean a canonical URL, product name, price as a decimal, currency code, availability enum, and evidence spans supporting each field. For page classification, define the allowed labels and how to handle uncertain or mixed pages.
- Target scope: allowed domains, paths, page types, and exclusions.
- Fields and types: required versus optional values, accepted formats, units, and null behavior.
- Update cadence: one-time dataset, periodic refresh, or near-current data.
- Success measures: field-level precision and recall, exact match, coverage, and an abstention rate for uncertain cases.
- Provenance: source URL, retrieval timestamp, response status, content hash, parser/model version, and evidence for each prediction.
Do not let the model silently invent missing values. Store a value as unknown or abstain when the page does not contain adequate evidence.
2. Choose an allowed data acquisition path
Check for an official API, feed, or export first. If HTML crawling is appropriate, review the site’s robots.txt, terms, licenses, privacy rules, authentication boundaries, and rate limits. OECD documents the increasing use of robots.txt and explicit terms restrictions for AI-training collection; those controls and applicable obligations should be checked before collection. See the OECD report on intellectual property issues in AI training. This engineering guide is not jurisdiction-specific legal advice.
For HTTP-accessible content, Scrapy provides requests, spiders, selectors, item pipelines, caching, and feed exports. Its spider model parses responses and yields items or further requests; pipelines and feed exports can then validate and persist records. See the Scrapy overview and official tutorial.
When a page is dynamic, first inspect the network requests and response data. If the information is available from an underlying request, reproducing that request is generally simpler and more reliable than rendering the whole page. Use Playwright when required data appears only after JavaScript execution, scrolling, or interaction. Scrapy’s dynamic-content guidance discusses these approaches; scrapy-playwright integrates browser rendering with Scrapy.
| Approach | Use it when | Tradeoffs to assess |
|---|---|---|
| Direct HTTP with Scrapy | Required information is in HTML or an underlying response. | Usually simpler to operate; depends on the response containing the data. |
| Scrapy with Playwright | Required information needs browser execution or interaction. | Can cover browser-dependent pages, with more compute, latency, and browser lifecycle complexity. |
| Hosted crawling/rendering or managed deployment | You need an operational service for rendering, crawl deployment, or monitoring. | Compare coverage, controls, observability, portability, rate-limit handling, and infrastructure cost. The Scrapy ecosystem lists options including Zyte API and Scrapy Cloud; verify current capabilities and terms directly. |
3. Create a Scrapy project and collect provenance
Install Scrapy in a virtual environment, create a project, and generate a spider. The following shell commands and spider are a runnable starting point for a site you are authorized to crawl. Replace the example domain and selectors with the target site’s documented structure. Start with a small allowed set of pages and a conservative request rate.
python -m venv .venv
# Linux or macOS:
source .venv/bin/activate
# Windows PowerShell:
# .venv\\Scripts\\Activate.ps1
python -m pip install Scrapy
scrapy startproject webdata
cd webdata
scrapy genspider catalog example.com
Replace webdata/spiders/catalog.py with this example. It keeps source metadata alongside extracted fields, follows product links on the same domain, and uses a deliberately simple selector example that you must adapt.
import scrapy
from datetime import datetime, timezone
from urllib.parse import urlparse
class CatalogSpider(scrapy.Spider):
name = "catalog"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
custom_settings = {
"ROBOTSTXT_OBEY": True,
"DOWNLOAD_DELAY": 1.0,
"CONCURRENT_REQUESTS_PER_DOMAIN": 2,
"FEEDS": {"products.jsonl": {"format": "jsonlines", "encoding": "utf8"}},
}
def parse(self, response):
for card in response.css("article.product-card"):
link = card.css("a.product-link::attr(href)").get()
if link:
yield response.follow(link, callback=self.parse_product)
# Adapt pagination selector to the permitted site.
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
def parse_product(self, response):
price_text = response.css("[itemprop='price']::attr(content)").get()
yield {
"url": response.url,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"http_status": response.status,
"domain": urlparse(response.url).hostname,
"name": response.css("h1::text").get(default="").strip() or None,
"price_raw": price_text,
"currency": response.css("[itemprop='priceCurrency']::attr(content)").get(),
"evidence": {
"name_selector": "h1",
"price_selector": "[itemprop='price'][content]",
},
}
Run it with scrapy crawl catalog; JSON Lines feed export is configured in the spider. Scrapy supports feed exports and storage backends, selectors, caching, cookies, authentication, and robots.txt handling; consult its feed export documentation for format and destination settings.
The example records provenance but not a raw response artifact. For audit and relabeling, retain an access-controlled raw HTML or response snapshot where permitted, together with retrieval time, status, URL, and content hash. Avoid storing secrets or unnecessary personal data. Use a deterministic canonical URL or content hash to detect duplicates, and make crawl limits explicit.
4. Render pages only when the data requires it
Inspect the page’s network activity and compare the initial HTML with the rendered page. If a JSON or HTML endpoint already contains the desired fields, prefer that response. If the page requires JavaScript or a user interaction, render only the relevant pages, bound browser concurrency, and close browser contexts predictably.
A minimal Playwright inspection script can help determine whether a field appears after rendering. Install it with python -m pip install playwright, then install the browser with playwright install chromium. Use this only on pages you are allowed to access.
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
response = await page.goto("https://example.com/product/1", wait_until="domcontentloaded", timeout=30000)
await page.locator("h1").wait_for(timeout=10000)
print({
"status": response.status if response else None,
"url": page.url,
"title": await page.locator("h1").inner_text(),
})
await browser.close()
asyncio.run(main())
This is an inspection example rather than a production crawler. In production, cap navigation and selector waits, record timeouts as explicit outcomes, and integrate browser work with queueing, retry, and provenance policies. For deeper integrations, use the scrapy-playwright project documentation.
5. Build labels and a baseline before training
Begin with a deterministic parser or rules baseline. Label representative pages by hand, including missing fields, unusual formats, new layouts, and ambiguous cases. Each training example should include the input evidence, expected structured output, source metadata, and label reviewer or review status. Keep the original text span or DOM evidence for every extracted field so reviewers can trace model predictions back to the page.
Use an LLM or smaller classifier where rules become brittle: classifying page type, extracting a field from varying layouts, normalizing a value, or suggesting duplicate matches. Treat model output as a prediction to validate, not as ground truth. Require a schema check, type check, allowed-value check, and evidence check before writing a record. Do not include secrets, private account pages, or personal data in training examples unless collection and processing are expressly permitted and appropriately governed.
6. Split data and train a task-sized model
Split by page and preferably by domain or time. Randomly splitting near-identical pages can put copies of the same template in both training and evaluation and make performance appear better than it is. Keep a test set untouched while tuning. If your target sites have distinct templates, reserve entire domains or template families for a harder generalization check.
Start with the simplest model that meets the measured need: selectors and normalization rules, a small page-type classifier, or a constrained extraction model. Fine-tuning is justified only when you have enough reviewed examples and a stable, measurable error pattern that prompting or deterministic logic does not address. Model development includes data preparation, training, evaluation, and improvement; see the OpenAI overview of model development for a general lifecycle description.
Version the dataset, schema, rules, model, and prompt together. Record model confidence where available, but calibrate its meaning against your validation data rather than treating a score as a guarantee.
7. Validate, deduplicate, and export
Validate required fields and types in a pipeline before export. Quarantine invalid records with their source and error details rather than silently dropping or coercing them. Deduplicate using a canonical URL, stable source identifier, or content hash according to the task; two pages with similar text are not necessarily the same record.
Scrapy’s item pipelines can validate and transform items, while feed exports can produce JSON, CSV, or JSON Lines. Preserve raw evidence and normalized data separately. JSON Lines works well for append-oriented processing; CSV can be convenient for tabular inspection but is less expressive for nested evidence fields.
8. Evaluate and monitor in production
Measure field-level precision and recall, exact-match rate for complete records, coverage, schema rejection rate, and abstention rate. Inspect false positives and false negatives by field, domain, layout, and failure cause. A single aggregate score can conceal a broken field or a failure limited to a new template.
In operation, monitor empty-field rates, validation failures, duplicate rates, crawl status codes, timeouts, browser latency, model latency, and shifts in page layout or value distributions. Log a sample of evidence and predictions with sensitive data minimized. Alert on sustained changes, and route low-confidence or novel layouts for review. The Scrapy project site lists ecosystem tools such as Spidermon for crawl validation and alerts; verify the current feature set before adopting it.
9. Performance, reliability, and cost
- Use direct requests first: browser startup and rendering add latency and resource use. Render only pages where measurements show that needed data depends on a browser.
- Control concurrency: tune per-domain concurrency and delay to respect site limits and avoid creating a retry storm. Apply bounded retries to transient network or server errors; do not retry permanent denials indefinitely.
- Cache carefully: caching reduces repeat requests during development and refreshes, but account for freshness requirements and site policies. Scrapy documents HTTP caching in its HTTP cache middleware documentation.
- Bound work: limit crawl depth, pages, response sizes, navigation timeouts, browser contexts, and queue growth. Record incomplete runs explicitly.
- Budget the full pipeline: include network transfer, browser compute, storage for raw evidence, human labeling and review, model inference or training, retries, and operations. Compare self-hosted and hosted approaches on these same units; the dossier contains no benchmark establishing one as universally cheaper or more accurate.
- Preserve portability: export normalized records in documented formats and retain schema/version metadata so a crawler or model can be replaced without losing the dataset.
10. Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Fields are empty in Scrapy responses | The content is injected by JavaScript, selectors changed, or the response differs from the browser view. | Inspect the response body and network requests; use the underlying endpoint if available, otherwise render only that page path. Add selector assertions and validation counts. |
| Browser waits time out | The chosen selector never appears, navigation is slow, or the site changed. | Check the selector and response status, wait for a specific required element rather than an arbitrary long delay, and record the timeout as a failed page outcome. |
| HTTP 403 or 429 responses | Access is denied or request volume exceeds allowed limits. | Stop aggressive retries, review terms and access requirements, reduce request rate, and use an authorized API or contact the site owner where appropriate. Do not attempt to bypass access controls. |
| Duplicate or contradictory records | URL variants, pagination overlap, or changed content are being treated as distinct or identical incorrectly. | Define canonicalization rules, retain retrieval timestamps and hashes, and deduplicate according to stable identifiers and task semantics. |
| High evaluation score but poor live results | Near-duplicate pages leaked across splits, or the test set missed new domains/layouts. | Split by domain or time, evaluate on new templates, and audit field-level errors and evidence spans. |
| Model returns invalid or unsupported values | Output was accepted without schema validation or the page lacks enough evidence. | Validate types and enums, preserve evidence, reject or abstain on unsupported predictions, and add reviewed examples for recurring error patterns. |
| Crawl appears stuck or memory grows | Unbounded link following, large responses, too much concurrency, or browser contexts not being closed. | Set depth/page limits, restrict allowed domains and paths, cap concurrency and response handling, and close browser resources in guaranteed cleanup paths. |
11. Or skip the browser setup
If the task is to capture pages as visual artifacts for review or downstream processing, ScreenshotNeo is a website screenshot API and MCP server. It can return PNG, JPEG, WebP, or PDF from one request; consult the ScreenshotNeo API documentation for parameters and response behavior.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the page verdict and billing status in headers.
- An MCP server exposes
take_screenshot,get_page_info, andcapture_pdffor Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000; every feature is on every plan.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
12. Frequently asked questions
Do I need to train a model from scratch?
Usually not. Begin with selectors, parsers, and a baseline classifier or extraction prompt. Consider fine-tuning only after reviewed data shows a persistent, measurable gap.
Is Playwright required for web scraping?
No. Use direct HTTP when the response or underlying request contains the required data. Add browser rendering for genuinely JavaScript-dependent or interactive content.
What should I save with each training example?
At minimum, retain the source URL, retrieval time, response status, schema and parser/model version, label status, and the evidence supporting each field, subject to permission and data-minimization requirements.
How do I know whether the model is ready?
Use a held-out evaluation that reflects new pages, layouts, domains, or time periods, and set explicit field-level quality and coverage thresholds before deployment. Continue monitoring those measures after release.


