How to Structure and Clean Web Data for AI
A practical workflow for selecting, extracting, deduplicating, validating, and refreshing web data before search or LLM ingestion.

Direct answer: clean web data for AI by defining the questions it must answer, selecting only relevant URLs, verifying crawler access, canonicalizing duplicates, extracting content with its meaning intact, storing consistent fields and provenance, validating against the source, and refreshing records as pages change. The best format is the one your destination accepts reliably: plain text, JSON, Markdown, HTML, PDF, or a task-specific schema. No markup file guarantees inclusion in an AI answer.
This workflow applies to retrieval-augmented generation (RAG), enterprise search, content recommendations, analytics, and public datasets. It separates durable data-quality work from product-specific ingestion rules.
1. Define the task and collection boundary
Start with the questions your AI workflow must answer. A support assistant may need product documentation and release notes; a research index may need articles, tables, and citations; a catalog assistant may need product records and availability dates. Write these requirements before writing a scraper.

- Specify entities and fields. For an article, this might be title, author, publication date, headings, body, cited URLs, and source URL.
- Choose URL patterns deliberately. Include canonical content paths and exclude account pages, faceted navigation, internal search results, print views, and tracking-parameter variants.
- Set an ownership rule. Record who is responsible for correcting a source, approving extracted data, and deciding when a page is removed.
- Define freshness. A changelog may need frequent refreshes; a historical report may be immutable. There is no universal schedule.
Google Cloud Agent Search recommends include and exclude URL patterns before indexing. Its documentation also warns that each unique URL is treated as a separate document, so uncontrolled variants can duplicate results and increase storage costs. Apply the same discipline even when your destination is a different search or vector system.
2. Make sure pages can be fetched and rendered
A perfect extractor cannot clean content that the ingestion crawler cannot access. Test representative URLs from each template and rendering mode.
- Check DNS, TLS, redirects, authentication, and firewall rules from the crawler’s network.
- Review
robots.txt, sitemap availability, and any vendor-specific crawler requirements. - Test pages whose main content is inserted by JavaScript. Google Search can process JavaScript when it is not blocked, but JavaScript-heavy sites require more careful technical work.
- Record HTTP status, final URL, retrieval time, and a short failure reason for every fetch.
Do not assume that a successful browser load means an ingestion service sees the same document. Google Cloud Agent Search, for example, uses its own crawler and separately fetches sitemaps with Googlebot; other systems have different behavior.
3. Canonicalize URLs and remove duplicates
Duplicate pages waste storage, confuse ranking, and make updates inconsistent. Normalize URLs before extraction and again before indexing.
Useful normalization rules
- Resolve relative links against the source URL.
- Lowercase the host and remove the default port.
- Remove fragments such as
#comments. - Drop known tracking parameters such as
utm_sourcewhen they do not change content. - Sort remaining query parameters only when the site treats their order as irrelevant.
- Follow redirects and retain the final canonical URL.
- Prefer the page’s
link rel="canonical"when it is valid and belongs to an allowed host.
Keep both the original URL and normalized URL for auditability. A hash of normalized, meaningful text can catch copies whose URLs differ. For near-duplicates, compare shingles or embeddings and send uncertain matches to review instead of deleting automatically.
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
TRACKING = {"utm_source", "utm_medium", "utm_campaign", "utm_term", "utm_content", "gclid"}
def canonicalize(raw: str) -> str:
p = urlsplit(raw.strip())
query = [(k, v) for k, v in parse_qsl(p.query, keep_blank_values=True)
if k.lower() not in TRACKING]
return urlunsplit((p.scheme.lower(), p.netloc.lower(), p.path or "/", urlencode(sorted(query)), ""))
4. Extract content without destroying meaning
Cleaning means removing noise while preserving information needed to answer the task. Keep headings in order, list boundaries, table headers, units, links, dates, and relationships between labels and values. Navigation, cookie notices, repeated footers, and recommendation rails can usually be removed when they are not part of the subject.
Semantic HTML improves human readability and accessibility, but Google Search Central says you do not need perfectly semantic or valid HTML for its systems to understand a page. Treat extraction as a transformation that must be checked against the original, not as proof that the result is correct.
A small, reproducible extraction script
import json
import hashlib
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
url = "https://example.com/article"
r = requests.get(url, timeout=30, headers={"User-Agent": "ResearchBot/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
for node in soup.select("script, style, nav, footer, aside, form, .cookie-banner"):
node.decompose()
main = soup.select_one("main, article") or soup.body
for heading in main.select("h1, h2, h3, h4"):
heading.name = heading.name # preserve heading levels
text = "\\n".join(line.strip() for line in main.get_text("\\n").splitlines() if line.strip())
record = {
"id": hashlib.sha256(url.encode()).hexdigest(),
"source_url": url,
"retrieved_at": "2026-09-29T00:00:00Z",
"title": (soup.find("h1").get_text(" ", strip=True) if soup.find("h1") else None),
"text": text,
"content_sha256": hashlib.sha256(text.encode()).hexdigest(),
}
print(json.dumps(record, ensure_ascii=False))
Production extraction should handle missing main elements, malformed HTML, language detection, tables, embedded JSON, and pages that require rendering. Preserve the raw response or an immutable snapshot when licensing and policy permit; it gives reviewers something to compare with the cleaned record.
5. Choose a consistent representation
Use stable field names, predictable types, and identifiers that survive re-fetches. A practical JSON record might contain:
{
"id": "stable-source-record-id",
"source_url": "https://example.com/page",
"canonical_url": "https://example.com/page",
"retrieved_at": "2026-09-29T00:00:00Z",
"title": "Page title",
"language": "en",
"content": "Clean text with headings preserved",
"links": [{"url": "https://example.com/ref", "label": "Reference"}],
"metadata": {"section": "docs", "published_at": "2026-01-19"},
"content_sha256": "..."
}
JSON-LD can map terms to IRIs through a shared context, making variable documents more deterministic across systems. It is an option, not a universal requirement. Destination systems may accept TXT, JSON, Markdown, HTML, PDF, DOCX, PPTX, XLSX, or XLSM; Google Cloud Agent Search lists these formats in its unstructured-data ingestion guidance.
For JSON Lines output, validate one object per line and reject records with missing identifiers. For Markdown, use one heading hierarchy and explicit link destinations. For plain text, include labels before values so a model can distinguish a price from a date or a person’s name.
6. Validate accuracy, provenance, and safety
Run automated checks before indexing:
- Required fields exist and have the expected type.
- URLs use permitted schemes and hosts.
- Dates parse consistently and are not silently shifted by timezone conversion.
- Text is not empty, truncated, or dominated by navigation boilerplate.
- Canonical URLs are unique within the collection.
- Extracted values match the source snapshot or response.
- Secrets, personal data, and restricted fields are removed or access-controlled.
Sample records should receive human review, especially when tables, legal language, medical information, prices, or identity data are involved. The UK government’s AI-ready data framework emphasizes quality, metadata, APIs, governance, stewardship, and human-in-the-loop checks. Keep an audit trail containing source URL, retrieval date, extractor version, validation result, and reviewer decision.
7. Refresh and monitor the collection
Store a content hash and compare it on each fetch. Re-index only changed records when the destination supports incremental updates. Monitor:
- new, removed, redirected, and repeatedly failing URLs;
- changes in word count or heading structure;
- duplicate rates after each crawl;
- schema-validation failures;
- stale records beyond your freshness target.
Choose the refresh interval from the source’s change rate and the cost of stale answers. A failed fetch should normally preserve the last known record with a visible stale status instead of deleting good data immediately.
8. Do you need special schema markup for AI search?
For Google’s generative AI search features, publicly accessible, crawlable pages and established technical practices remain central. Google Search Central states: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add.” Continue using structured data when it accurately describes a page and supports eligible search features, then validate it against the applicable guidelines.
Do not treat an AI-specific manifest as a guarantee of citation or visibility. LLM-LD is a draft proposal from CAPXEL that describes crawl-ready, ingest-ready, and agent-ready levels and files such as llm-index.json; it is not a general requirement or an independently established standard.
9. Capture rendered pages for your pipeline
Some cleaning workflows need a reliable rendered snapshot before text extraction: client-side documentation, authenticated previews, or pages whose consent overlays obscure the article. You can run a browser yourself with Playwright or Puppeteer, wait for a selector or network idle, hide overlays, and save HTML or a screenshot. This gives control but adds browser binaries, concurrency limits, retries, and maintenance for changing sites.
10. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It can load lazy images, capture an element by CSS selector, set a viewport or device preset, apply custom CSS and JavaScript, wait for a selector, delay, or network idle, and provide headers, cookies, user agents, authorization, timezone, and geolocation. Use the ScreenshotNeo API documentation for all options.

cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts 63 options, including full-page capture, dark mode, retina scale, PDF paper size and page ranges, resizing, caching with a chosen TTL, blocking ads, trackers, requests or resource types, click actions, hidden selectors, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can reduce migration changes.
Before capture it accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Every response identifies the page verdict and billing result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
There is a free allowance of 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account to try the capture step before adding it to your ingestion pipeline.
11. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| Many empty records | Content is rendered after initial HTML or blocked by a firewall. | Use a rendering-capable fetcher, wait for a meaningful selector, and verify crawler access. |
| Duplicate search results | Tracking, faceted, print, or locale URL variants were indexed. | Normalize query parameters, honor canonical URLs, and define include/exclude patterns. |
| Tables lose meaning | Cells were flattened without headers or units. | Emit headers with each row and preserve units, captions, and row labels. |
| Fresh pages look stale | Cache or change detection is too aggressive. | Record retrieval time, compare content hashes, and shorten the refresh interval. |
| 403 or 429 responses | Rate limits, authentication, or bot defenses. | Use authorized access, backoff and retry, lower concurrency, and follow site policy. |
| Wrong dates or prices | Locale or timezone conversion changed values. | Store the original string plus normalized value and timezone; review samples. |
| Screenshot contains a popup | Overlay appeared after the first page load. | Wait for the overlay selector, click consent, or hide it before capture. |
12. Performance, reliability, and cost
Measure the whole pipeline: fetch latency, render time, extraction time, duplicate rate, validation failures, and destination indexing latency. Limit concurrency to what the source and your browser workers can sustain. Use exponential backoff for transient errors, idempotent record IDs, and a dead-letter queue for pages requiring review.
Cache immutable or unchanged pages, but attach a TTL and retrieval timestamp. Batch operations where supported; ScreenshotNeo bulk capture handles up to 100 URLs per call. Avoid paying to process duplicates by canonicalizing before rendering. Keep raw and cleaned data in separate stores so an extractor change can be replayed without fetching every source again.
FAQ
What format should web data be in for an LLM?
Use the format your retrieval or ingestion system accepts reliably. JSON or JSONL is useful for typed fields and provenance; Markdown or text works well for readable passages; HTML or PDF may preserve document structure. Consistency and traceability matter more than a fashionable extension.
How do I remove duplicate pages before indexing?
Canonicalize URLs, remove non-content query parameters, follow redirects, honor valid canonical links, and compare content hashes. Review near-duplicate clusters before deleting records.
Does AI search need special schema markup?
Google says special schema.org markup is not required for generative AI search. Keep accurate structured data for supported uses, validate it, and focus on crawlability, useful content, and low duplication.
Should every extracted value be reviewed by a person?
No. Automate syntax, type, URL, and consistency checks, then route high-impact fields and uncertain transformations to human review with source evidence attached.


