ScreenshotNeo

BlogEngineering

AI Training Data Collection with Web Crawlers

Learn how AI crawlers discover, fetch, filter, and govern web data—and how publishers can control training and search access.

By the ScreenshotNeo team1 October 202610 min read

AI companies collect training data through a pipeline: discover URLs, fetch pages, record response metadata, extract and normalize content, apply permission and quality filters, deduplicate records, and store each item with provenance. A page being publicly reachable does not by itself settle copyright, privacy, terms-of-service, or crawler-policy questions.

This guide explains how the pipeline works, what robots.txt can and cannot do, how search and training crawler controls differ, how Common Crawl data is structured, and what publishers can do to manage access.

How web crawling becomes training data

A production collection system usually has these stages:

  1. Discovery: seed URLs come from submitted lists, links found on other pages, sitemaps, public datasets, or prior crawls. A URL frontier queues work and prevents duplicate requests.
  2. Permission checks: the crawler identifies its user agent, retrieves robots.txt, evaluates the applicable rule, and applies internal allowlists, blocklists, opt-outs, terms, and licensing policies.
  3. Fetching: workers request pages while respecting rate limits, retry rules, TLS validation, redirect limits, and response-size limits. They record status code, headers, final URL, retrieval time, and failure reason.
  4. Parsing: HTML is decoded, boilerplate is separated from main content, links and metadata are extracted, and structured formats such as JSON-LD may be retained.
  5. Normalization: canonical URLs, character encoding, whitespace, language labels, timestamps, and content types are standardized.
  6. Filtering: spam, malware, adult material, duplicates, navigation-only pages, and selected unwanted personal-data sources can be removed. OpenAI describes filtering publicly available sources in this way, but implementations differ by provider.
  7. Dataset construction: records are joined with provenance, policy decisions, hashes, license signals, and crawl version identifiers. Training or evaluation datasets are then built from the approved records.

There is no single universal implementation. A crawler may use a browser for JavaScript-heavy pages, while another uses HTTP fetches only. Vendor-specific behavior should be documented instead of assumed.

Does robots.txt stop AI training crawlers?

robots.txt is an operational instruction that compliant crawlers download and parse before crawling. Google documents that crawlers select the most specific matching user-agent group. A rule can therefore affect whether a particular bot requests a URL, but it is not a complete legal license, copyright waiver, or guarantee that every automated client will obey it.

Publishers should treat robots rules as one control in a broader governance record that also includes terms of service, licenses, privacy obligations, consent or opt-out signals, takedown procedures, and evidence of when a policy changed.

Example policy that blocks a named training bot while allowing ordinary search crawlers:

User-agent: GPTBot
Disallow: /

User-agent: *
Allow: /

Google’s crawler documentation explains user-agent matching and robots parsing: Robots.txt specifications.

Yes, when the relevant products use separate crawlers and controls. OpenAI documents GPTBot as the crawler associated with content that may be used to train foundation models and OAI-SearchBot as the crawler used for search presentation. Each setting is independent of the others, so a publisher can make a search-visibility decision separately from a training-use decision.

For example:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: *
Allow: /

OpenAI notes that robots changes can take about 24 hours to affect search crawling behavior. Keep a dated change log and verify requests in your server logs after changing a group. See OpenAI’s crawler overview for the current bot names and controls.

What Common Crawl provides

Common Crawl describes a corpus with three useful layers:

  • Raw page data: captured web responses for researchers who need the original bytes and headers.
  • Metadata extracts: crawl-level and response-level fields that help locate, filter, and audit records.
  • Text extracts: processed text that is easier to search or feed into downstream pipelines.

Common Crawl’s terms permit use in connection with AI systems, including developing, training, or deploying them. The terms also warn that crawled pages can carry separate third-party rights and that users remain responsible for applicable law. Treat a Common Crawl record as a source artifact, not as proof that every downstream use is licensed.

Read the Common Crawl terms of use and its corpus overview before building a dataset.

Build a respectful crawler: runnable examples

Python: fetch robots.txt, crawl a page, and retain provenance

This small example uses Python’s standard library. It checks robots permissions, identifies the crawler, follows one page, and records enough metadata to audit the fetch. It is intentionally conservative: one URL, a delay, a size limit, and no JavaScript execution.

#!/usr/bin/env python3
import json
import time
from datetime import datetime, timezone
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
from urllib.request import Request, urlopen

USER_AGENT = "ExampleResearchBot/1.0 (+https://example.org/bot-info)"
MAX_BYTES = 2_000_000
DELAY_SECONDS = 2

def fetch(url: str) -> dict:
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    rp = RobotFileParser(robots_url)
    try:
        rp.read()
        allowed = rp.can_fetch(USER_AGENT, url)
    except Exception as exc:
        return {"url": url, "decision": "error", "reason": f"robots fetch failed: {exc}"}

    if not allowed:
        return {"url": url, "decision": "blocked_by_robots", "retrieved_at": datetime.now(timezone.utc).isoformat()}

    time.sleep(DELAY_SECONDS)
    request = Request(url, headers={"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
    try:
        with urlopen(request, timeout=30) as response:
            body = response.read(MAX_BYTES + 1)
            truncated = len(body) > MAX_BYTES
            body = body[:MAX_BYTES]
            return {
                "url": url,
                "final_url": response.geturl(),
                "status": response.status,
                "content_type": response.headers.get("Content-Type"),
                "etag": response.headers.get("ETag"),
                "last_modified": response.headers.get("Last-Modified"),
                "retrieved_at": datetime.now(timezone.utc).isoformat(),
                "truncated": truncated,
                "body_utf8": body.decode("utf-8", errors="replace")
            }
    except Exception as exc:
        return {"url": url, "decision": "fetch_error", "error": str(exc)}

if __name__ == "__main__":
    result = fetch("https://example.com/")
    print(json.dumps(result, indent=2))

Node.js: the same controlled fetch

const { URL } = require('node:url');

const userAgent = 'ExampleResearchBot/1.0 (+https://example.org/bot-info)';

async function fetchPage(target) {
  const url = new URL(target);
  const robots = await fetch(new URL('/robots.txt', url), {
    headers: { 'User-Agent': userAgent }
  });
  const robotsText = robots.ok ? await robots.text() : '';
  if (/^Disallow:\s*\/\s*$/mi.test(robotsText)) {
    return { url: target, decision: 'blocked_by_robots' };
  }

  await new Promise(resolve => setTimeout(resolve, 2000));
  const response = await fetch(target, {
    redirect: 'follow',
    headers: {
      'User-Agent': userAgent,
      'Accept': 'text/html,application/xhtml+xml'
    },
    signal: AbortSignal.timeout(30000)
  });
  const body = (await response.text()).slice(0, 2_000_000);
  return {
    url: target,
    finalUrl: response.url,
    status: response.status,
    contentType: response.headers.get('content-type'),
    retrievedAt: new Date().toISOString(),
    body
  };
}

fetchPage('https://example.com/').then(result => console.log(JSON.stringify(result, null, 2)));

The Node example demonstrates the flow, but a production parser should implement the robots specification fully, including user-agent groups, wildcard matching, and crawl-delay policy where applicable.

cURL: inspect a policy and a response

curl --fail --max-time 30 -A 'ExampleResearchBot/1.0 (+https://example.org/bot-info)' \
  https://example.com/robots.txt

curl --fail --max-time 30 -D response-headers.txt \
  -A 'ExampleResearchBot/1.0 (+https://example.org/bot-info)' \
  -o page.html https://example.com/

Filtering, deduplication, and personal-data handling

A useful dataset records why every item was retained or rejected. Common controls include:

Control What to record Failure to avoid
Canonicalization Original URL, canonical URL, redirect chain Counting tracking-parameter variants as separate pages
Deduplication Exact hash and near-duplicate fingerprint Overweighting syndicated or mirrored text
Language Detected language and confidence Mislabeling multilingual pages
Quality Spam, boilerplate, malware, and thin-content decisions Letting navigation or error pages enter training data
Personal data Detection, minimization, redaction, retention decision Publishing unnecessary contact or account information
Provenance Crawl ID, timestamp, source, policy evidence, license signal Being unable to trace or remove a record later

Filtering decisions should be reproducible. Store the rule version, model or classifier version, and reason code alongside each record. Keep raw material access-controlled when retention is necessary for auditing.

How to compare AI data-collection approaches

No single corpus or crawler wins on every dimension. Evaluate these axes before selecting a source:

  1. Permission and opt-out handling: Does the system parse robots groups, honor explicit opt-outs, and preserve evidence?
  2. Coverage: Which sources, languages, domains, and geographic regions are represented?
  3. Freshness: How often are pages recrawled, and are update times retained?
  4. Filtering: How are spam, malware, adult content, boilerplate, and duplicates handled?
  5. Personal-data minimization: Are sensitive fields removed before release or training?
  6. Provenance: Can each training item be traced to a URL, retrieval time, and crawl version?
  7. Licensing: What downstream-use terms apply, and how are third-party rights surfaced?
  8. Infrastructure behavior: What are the rate limits, retry rules, storage costs, and failure modes?

Publisher checklist for AI crawler governance

  1. Inventory crawler user-agent strings and classify them as search, training, advertising, monitoring, or user-triggered access.
  2. Write separate robots.txt groups for the decisions you actually want to make.
  3. Publish terms of service and licensing language that matches your intended reuse rules.
  4. Record consent, opt-out, takedown, and personal-data handling procedures.
  5. Keep dated copies of policy files and test them against representative user-agent strings.
  6. Log request time, user agent, URL, status, response size, and policy decision.
  7. Require provenance, deduplication, and personal-data checks before releasing a dataset.
  8. Review policies regularly because crawler behavior, standards interpretation, and AI copyright rules change.

Or skip the browser setup

If your workflow needs screenshots of pages for visual datasets, documentation, or agent tasks, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in headers.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', image);

ScreenshotNeo includes full-page and element captures, device presets, custom viewports, retina scale, CSS and JavaScript, selectors to hide, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture, usage data, and PDF output. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for AI clients.

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Troubleshooting

My crawler receives 403 or 429 responses

Cause: the site is rate-limiting, blocking your network, or requiring authentication. Fix: slow the request rate, identify your user agent with a contact URL, honor Retry-After, stop on repeated denials, and obtain permission where required.

The robots result is ambiguous

Cause: malformed syntax, multiple groups, redirects, or an unavailable robots file. Fix: retain the raw file and retrieval status, apply the parser’s documented fallback, and use a conservative policy when you cannot establish permission.

Pages are empty or mostly navigation

Cause: content is rendered by JavaScript, blocked by a consent wall, or extracted with a poor main-content rule. Fix: classify the response, record the failure reason, and use a browser renderer only when policy and resource limits allow it. Do not silently treat an empty result as valid training text.

Duplicate records dominate the dataset

Cause: URL parameters, mirrors, syndicated articles, or repeated crawl snapshots. Fix: normalize URLs, use exact and near-duplicate hashes, and retain one canonical record while preserving provenance links.

A publisher requests removal

Fix: locate all records through URL and content hashes, preserve the request and decision, remove or quarantine affected derivatives, and document the completion date. A robots change alone may not identify data already collected.

Performance, reliability, and cost

  • Throughput: concurrency increases load and blocking risk. Use per-host queues, bounded workers, connection reuse, and adaptive delays.
  • Reliability: separate transient failures from permanent policy denials. Retry timeouts and 5xx responses with capped exponential backoff; do not retry a robots denial as if it were a network error.
  • Freshness: use conditional requests with ETag or Last-Modified when supported, and schedule recrawls by change frequency.
  • Storage: compress raw responses, store text and metadata separately, and retain hashes so unchanged pages do not create full duplicate copies.
  • Cost: budget DNS, bandwidth, browser compute, parsing, storage, filtering, and legal review. Browser rendering is usually more expensive than HTTP fetching, so reserve it for pages that need it.
  • Auditability: keep policy snapshots, request logs, classifier versions, and dataset manifests. These reduce the cost of responding to corrections or takedown requests.

FAQ

Is a public webpage automatically free to use for model training?

No. Public reachability does not resolve copyright, privacy, contract, licensing, or crawler-policy questions. The U.S. Copyright Office’s AI initiative is examining these issues, and outcomes depend on jurisdiction and facts.

Does allowing a search bot allow a training bot?

Not necessarily. Separate user-agent groups can express separate decisions, such as allowing OAI-SearchBot while disallowing GPTBot.

Can robots.txt remove data that was already crawled?

Usually it only controls future requests. Use a documented takedown process to find and remove existing records and derivatives.

What should every dataset record include?

At minimum: source URL, final URL, retrieval time, response metadata, content hash, crawl version, policy evidence, filtering decisions, and licensing or opt-out signals.

Is Common Crawl a licensed waiver for every page?

No. Its terms allow AI-related use while preserving responsibility for third-party rights and applicable law.