ScreenshotNeo

BlogGuides

How to Improve AI Models with Web Scraping

Learn how to collect, evaluate and govern web data so it improves a specific AI task instead of adding noisy pages.

By the ScreenshotNeo team1 October 20268 min read

Web scraping can improve an AI model only when the collected data fits a defined task and is evaluated, cleaned, documented and used lawfully. More pages do not automatically produce a better model. Start with the user need, choose sources that contain the required features, measure data and model quality, and record how the corpus was gathered and processed.

This guide presents a practical workflow, runnable collection code, quality checks, access and reuse controls, Common Crawl options, performance guidance and failure fixes.

1. Define the task before collecting pages

Write down four things before choosing a crawler or corpus:

  1. Task: What should the model do—classify, retrieve, summarize, answer questions, extract fields or generate text?
  2. Users: Who will rely on the output, and what errors are costly?
  3. Required coverage: Which languages, domains, document types, dates and edge cases must appear?
  4. Success measure: Which evaluation set and user-facing metric will show improvement?

Web data is one possible input to model development. Teams may also use partner material, human-provided information and generated data. The role of data differs across preparation, pre-training, post-training and later evaluation. OpenAI describes these stages and data sources.

2. Choose an existing corpus or collect purpose-built data

Option Use it when Questions to answer
Existing corpus You need broad experimentation quickly. Does it contain the domains, languages, dates and fields your task needs? Can you process and govern it?
Purpose-built crawl You need current, narrow or structured pages. Can you access the sources under their controls and terms? How will you handle failures, duplicates and updates?
Hybrid A broad base needs targeted gaps filled. Can you document provenance and keep evaluation slices comparable?

Google PAIR recommends evaluating breadth, features, quality and collection methods, then documenting the dataset and gathering and processing decisions in a data card or equivalent record. See Data Collection + Evaluation.

3. Use a respectful collection pipeline

A reliable pipeline separates discovery, fetching, parsing, filtering, storage and evaluation. Keep the original URL, retrieval time, response status, content type, parser version and processing decisions with every record.

Minimal Python collector

from urllib.parse import urljoin
import json
import time
import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/"
HEADERS = {"User-Agent": "ResearchBot/1.0 (contact: data-team@example.org)"}

session = requests.Session()
session.headers.update(HEADERS)

response = session.get(START_URL, timeout=30)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
for tag in soup(["script", "style", "noscript"]):
    tag.decompose()

record = {
    "url": response.url,
    "retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
    "status": response.status_code,
    "content_type": response.headers.get("content-type", ""),
    "title": soup.title.get_text(" ", strip=True) if soup.title else "",
    "text": soup.get_text(" ", strip=True),
    "links": [urljoin(response.url, a["href"]) for a in soup.select("a[href]")],
}

with open("record.jsonl", "w", encoding="utf-8") as output:
    output.write(json.dumps(record, ensure_ascii=False) + "\n")

For a real crawl, add a queue, a per-host rate limit, retry limits, content-size limits, persistent checkpoints and a parser for each document type. Do not treat this small example as permission to crawl a site at scale.

Respect crawler controls and source terms

  • Read the site’s robots.txt and robots meta directives before fetching.
  • Check the source’s terms, licenses and any owner-specific opt-out mechanism.
  • Remember that crawler controls are service-specific. Google documents robots controls and Google-Extended for future Gemini training in its crawling guidance.
  • Keep a denylist for hosts or paths that must not be collected, and stop when an owner requests removal.

A public page is not automatically open data for unrestricted reuse. Privacy, intellectual-property, cybersecurity and data-governance issues can apply. Get advice appropriate to your jurisdiction and project.

4. Parse and normalize records

Store raw responses separately from normalized text so you can reprocess a parser without recrawling. A useful normalized record contains:

  • Canonical URL and redirect chain
  • Retrieval timestamp and HTTP metadata
  • Language and document type
  • Title, headings, body text and structured fields
  • Hash of raw content and normalized content
  • Source permissions, exclusions and removal status
  • Parser and cleaning version

Remove navigation, repeated boilerplate and accidental markup only with rules you can inspect. Keep a sample of discarded text so an evaluator can detect over-cleaning.

5. Improve quality before model training

Use a staged quality review rather than assuming that a large crawl is useful.

  1. Relevance: Sample records against the task definition. Remove pages that cannot contribute required features.
  2. Validity: Reject truncated responses, error pages, empty documents and content that is not the expected type.
  3. Language and format: Detect unexpected languages, encodings and file types.
  4. Duplication: Measure exact and near-duplicate rates. Keep the rule and threshold in your dataset documentation.
  5. Contamination: Prevent evaluation and test material from leaking into training data.
  6. Safety and privacy: Apply project-specific filters and escalation paths for personal or sensitive information.
  7. Human review: Inspect random samples and difficult slices; record disagreements and revise the policy.

There is no universal deduplication or filtering recipe. The right policy depends on the task, users and risk tolerance. Compare a baseline model with and without each major data change on a fixed evaluation set.

6. Use Common Crawl for broad experiments

Common Crawl provides raw page data, metadata extracts and text extracts. Its corpus is hosted on AWS public datasets and can be analyzed there or downloaded. This makes it a practical experimentation route, but its breadth does not guarantee task fit.

Common Crawl’s homepage reports more than 300 billion pages spanning 15 years and 3–5 billion new pages each month; these are provider-reported, volatile figures. Its Terms of Use state that it cannot guarantee the truthfulness, authenticity, quality, lawfulness or accuracy of crawled content. Treat every record as an input to validation, not as ground truth, and check source-owner terms before reuse.

7. Evaluate whether the model actually improved

Run the same evaluation before and after the data change. Report results by the slices that matter to users: domain, language, document age, input length and difficult examples. Also track regressions, refusal or safety behavior where relevant, latency and operating cost.

If results improve, preserve the exact corpus version, filters, prompts or training configuration and evaluation outputs. If results do not improve, inspect task fit and data quality before collecting more pages. Scraping alone is not evidence of a better model.

8. Performance, reliability and cost

Fetching

  • Use connection pooling and bounded concurrency, but enforce per-host pacing.
  • Set connect and read timeouts separately when your HTTP client supports them.
  • Retry transient network failures with exponential backoff; do not retry permanent authorization or policy errors indefinitely.
  • Cache responses by URL and relevant request headers. Record cache age so stale data is visible.

Storage and processing

  • Compress raw responses and normalized text.
  • Process in partitions so one malformed document does not stop a run.
  • Checkpoint queue state and write idempotent outputs keyed by a content hash.
  • Estimate cost from fetched bytes, storage, parsing compute and model-evaluation runs before scaling.

Reliability

  • Capture status counts, timeout counts, parser failures, duplicate rates and records removed by policy.
  • Keep a dead-letter queue for pages requiring manual review.
  • Version robots and terms decisions because access conditions can change.

9. Troubleshooting common failures

Symptom Likely cause Fix
Many 403 or 429 responses Rate too high, blocked user agent or missing permission. Slow down, identify the crawler, follow the site’s controls and stop where access is not allowed.
Pages contain only a shell Content is rendered by JavaScript. Use an authorized browser-rendering workflow, or find an available structured source. Record that rendering was required.
Text is mostly menus or cookie notices Parser selected the whole document. Use semantic content selectors, remove repeated boilerplate and validate with sampled pages.
Dataset is large but model quality is flat Low task fit, duplicates or noisy sources. Measure relevance and duplicate rates, add targeted examples and compare on fixed slices.
Training and evaluation scores look unusually high Evaluation contamination or near duplicates. Hash and compare splits, separate collection windows and inspect overlapping examples.
Records disappear after a rerun Source changed, parser changed or a policy exclusion was applied. Version raw data, parser rules and exclusion decisions; compare manifests between runs.
Legal or privacy review blocks release Unclear rights, sensitive data or missing provenance. Pause reuse, document source terms and processing, remove or protect affected data, and obtain appropriate advice.

10. Capture rendered pages when visual layout is part of the data

Some tasks need a page image—for example, evaluating visual layout or supplying screenshots to a multimodal pipeline. A browser-based capture must handle JavaScript, lazy images, consent banners, popups, bot checks and timeouts. Keep screenshots tied to the same URL, timestamp and provenance record as extracted text.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

See the ScreenshotNeo API documentation for the full option list. The same API supports full-page screenshots with lazy images loaded, CSS element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/article \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', image);

An MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does scraping more pages always improve a model?

No. Improvement depends on task fit, coverage, quality, processing and evaluation. More noisy or duplicated pages can add cost without improving the target metric.

Is a publicly reachable page free to reuse?

Not automatically. Check crawler controls, source terms, privacy, intellectual-property and governance requirements for your project.

Should I start with Common Crawl?

Use it for broad experimentation when its coverage fits your task. Validate records carefully and review its terms and source-owner conditions.

How do I prove the data helped?

Compare a baseline and changed model on the same held-out evaluation set, report meaningful slices and preserve the corpus and processing versions.

When are screenshots useful?

Use them when visual layout or rendered content is part of the task. Keep screenshot provenance aligned with the corresponding URL and retrieval metadata.