ScreenshotNeo

BlogGuides

Web Scraping for RAG: How to Collect and Prepare Website Content

Build a reliable website-to-RAG pipeline: discover pages, check access, extract and clean content, chunk it for retrieval, and keep it fresh.

By the ScreenshotNeo team4 October 202612 min read

To prepare website content for retrieval-augmented generation (RAG), build an ingestion pipeline that discovers in-scope URLs, respects site access instructions, fetches pages, extracts their main content, normalizes and deduplicates it, preserves provenance, chunks and embeds useful passages, and refreshes them when sources change. Treat a sitemap as a discovery aid, not permission to access a page; treat robots.txt as crawler guidance, not confidentiality or access control.

A useful RAG corpus is more than a pile of scraped HTML. Every passage should retain enough structure to make sense on its own and enough metadata to trace it back to the source page and retrieval time.

1. Define scope and access before collecting

Write down which site, paths, page types, and languages belong in the corpus, what the content will be used for, and which crawler identity will make requests. Check the site’s crawler preferences, terms, authentication boundaries, and request limits. Do not bypass login requirements, CAPTCHAs, or other access restrictions.

robots.txt communicates how site owners prefer crawlers to interact with URLs. Google explains that it is not a way to keep a page out of search results; password protection and noindex are examples of different mechanisms for those purposes. A sitemap likewise does not grant access. Verify that the crawler identity you use can fetch both the sitemap and the pages you plan to ingest. [Google crawling guidance; Google’s robots.txt guide]

Set operational limits

  • Set an explicit URL and path scope; do not let link discovery expand without bounds.
  • Use a descriptive user agent and contact details where appropriate.
  • Choose a conservative request rate and concurrency, and honor server responses and retry guidance.
  • Set timeouts and maximum response sizes. Reject unexpected schemes and avoid fetching arbitrary user-supplied URLs from a privileged network, which can expose internal services.
  • Decide how authentication is supplied and stored. Keep credentials out of logs and indexed text.

Start with a sitemap when the site provides one, then combine it with a deliberately bounded seed list or links found on already in-scope pages. A sitemap can reveal new or updated URLs, but it does not guarantee that a crawler can fetch or index them. Google describes sitemaps as a way to help discover URLs and uses sitemap signals in crawling workflows; actual behavior still depends on the crawler and site configuration. [Google crawling guidance; Google Cloud data preparation]

Keep discovery bounded

  • Allow only the intended hostnames and path prefixes.
  • Skip fragments, logout links, search result pages, calendars, and other URL patterns that can generate unbounded variants.
  • Decide whether query parameters are meaningful. Strip tracking parameters such as campaign tags only when they do not change page content.
  • Keep a queue and a visited set keyed by your normalized URL so redirects and duplicate links do not trigger repeated work.
  • Record why each URL entered the queue, such as sitemap, seed list, or in-scope link. This helps explain unexpected corpus entries.

3. Fetch pages and normalize URL variants

For each fetch, record the requested URL, final URL after redirects, fetch timestamp, status code, content type, and relevant validators such as ETag or Last-Modified when present. Retain the original requested URL for debugging and provenance, and use the final or canonical URL as the deduplication key according to your policy.

Sites can expose the same content through trailing-slash variants, uppercase paths, tracking parameters, redirects, or alternate hostnames. Google Cloud’s website ingestion guidance recommends canonical URL handling to reduce duplicate URL variants. Canonicalization needs care: do not collapse URLs when the query or path changes the actual content, language, tenant, or access context. [Google Cloud: Prepare data for ingesting]

Practical URL handling

  1. Parse and validate the URL; allow only http and https and approved hosts.
  2. Resolve redirects within the allowed host policy and record the redirect chain or final URL.
  3. Remove known tracking parameters and normalize path and host consistently.
  4. Use a page’s canonical link as a signal, not an unquestioned command. Validate that it points to an in-scope equivalent page.
  5. Store both the normalized identity and the source URL so every indexed passage can be traced.

4. Extract main content and preserve its structure

Parse HTML into content rather than embedding raw source. Drop scripts, styles, repeated navigation, cookie notices, and boilerplate when they do not carry information needed by the application. Preserve headings, lists, tables, code samples, captions, and meaningful link labels; flattening these structures can make a passage ambiguous.

For example, a table row containing a value may be meaningless without its column headers. A paragraph under “Limitations” may need that heading attached to the extracted text. Google Cloud documents layout parsing and content-aware chunking for structured documents; this supports preserving layout context when it affects meaning. [Google Cloud: Parse and chunk documents]

JavaScript-rendered pages and other formats

A basic HTTP fetch sees the server response, which may contain the full article, a shell that expects JavaScript, or an access challenge. Inspect a sample of your target corpus before choosing a fetch method. If required content only appears after rendering, use a permitted browser-based capture or the site’s supported API. Handle PDFs and other document types with format-aware parsers, and verify that text extraction preserves page order, headings, and tables. There is no universally best parser for every corpus.

ScreenshotNeo can capture a rendered page as an image or PDF, but image or PDF captures are not a replacement for semantic HTML extraction when you need searchable text. For visual page records, audits, or a separate screenshot artifact, its API is one GET request and supports browser options such as waiting for a selector. See the ScreenshotNeo website and API documentation.

5. Clean content and attach provenance

Normalize character encoding and whitespace, remove repeated boilerplate, and detect empty or low-value pages before they reach the embedding step. Keep content transformations deterministic where possible, so a re-fetch can be compared with the earlier version.

Attach useful metadata to each document or passage. A practical record includes:

  • source_url: the canonical or normalized page identity.
  • requested_url and final_url: useful for redirect diagnosis.
  • title and heading path: context for passages.
  • fetched_at: when the source was collected.
  • published_at or updated_at: when available and trustworthy.
  • content_hash: to detect unchanged content.
  • content_type, language, and access scope: where relevant to retrieval filtering.

This metadata set is implementation guidance rather than a prescribed standard. Google Cloud’s guidance supports canonicalization and refresh workflows; AWS describes document cleaning, formatting, and chunking as data preparation before embeddings are used. [Google Cloud data preparation; AWS: Understanding RAG]

6. Chunk pages for retrieval, then embed

Chunking splits long documents into pieces that can be retrieved independently. Choose boundaries that keep a passage coherent: usually start at heading or paragraph boundaries, keep table headers with their rows, and carry the page title and relevant heading path into each chunk’s context or metadata. Very large chunks can bury the answer; very small chunks can lose definitions, conditions, or surrounding explanation.

There is no single chunk size that is right for every embedding model, content type, and question. Start with the structure of the source, use the embedding model’s documented input limits, then evaluate retrieval on representative questions. Review both whether the right source page is returned and whether the passage contains enough context to answer. Google Cloud describes parsing and chunking approaches; AWS and GOV.UK describe document preparation, embeddings or vectorisation, and indexing as steps in RAG pipelines. [Google Cloud parsing and chunking; AWS RAG guidance; GOV.UK: RAG Systems]

A useful chunking checklist

  • Split at semantic boundaries before falling back to fixed token windows.
  • Include title and heading context, especially when chunks are indexed separately.
  • Preserve table headers, list labels, units, and code language or filename where available.
  • Use overlap only where it helps preserve context across boundaries; too much overlap adds duplicate material to the index.
  • Keep source identifiers and offsets or section paths so retrieved text can cite its origin.
  • Version the chunking rules and embedding model. Re-embed when either changes.

7. Index, refresh, and remove stale content

Write prepared chunks and their metadata to the retrieval index. Keep a stable document identity so a changed page replaces or supersedes its previous chunks rather than accumulating duplicates. Use sitemap last-modified hints, validators, scheduled revisits, or other change signals to decide what to fetch again. Treat these as hints and compare content hashes where practical.

When a page disappears, becomes inaccessible, or is removed from scope, decide whether to delete its chunks or mark them unavailable. Avoid leaving stale passages retrievable without a clear policy. Store ingestion status separately from document text so failed refreshes do not accidentally erase a previously good version.

8. A minimal runnable fetch-and-extract example

This Python example fetches one explicitly chosen page, checks the response, extracts readable text from common article structure, and writes a JSON record with basic provenance. It is a starting point for a bounded ingestion worker, not a production crawler: add host allowlists, access-policy checks, rate limiting, retries, and persistent deduplication before scaling it up.

from datetime import datetime, timezone
from urllib.parse import urlparse
import hashlib
import json
import requests
from bs4 import BeautifulSoup

url = "https://example.com/docs/"
allowed_hosts = {"example.com"}
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or parsed.hostname not in allowed_hosts:
    raise ValueError("URL is outside the allowed scope")

response = requests.get(
    url,
    headers={"User-Agent": "ExampleRAGBot/1.0 (+https://example.com/contact)"},
    timeout=(5, 30),
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "text/html" not in content_type.lower():
    raise ValueError(f"Expected HTML, got {content_type!r}")

soup = BeautifulSoup(response.text, "html.parser")
for node in soup.select("script, style, noscript, nav, footer, header, aside"):
    node.decompose()
main = soup.find("main") or soup.find("article") or soup.body
if main is None:
    raise ValueError("No page content found")

title = soup.title.get_text(" ", strip=True) if soup.title else ""
text = "\n".join(line.strip() for line in main.get_text("\n").splitlines() if line.strip())
if len(text) < 100:
    raise ValueError("Page appears empty or too short to ingest")
canonical_tag = soup.find("link", rel="canonical")
canonical = canonical_tag.get("href") if canonical_tag else response.url
record = {
    "source_url": canonical,
    "requested_url": url,
    "final_url": response.url,
    "title": title,
    "fetched_at": datetime.now(timezone.utc).isoformat(),
    "status": response.status_code,
    "content_hash": hashlib.sha256(text.encode("utf-8")).hexdigest(),
    "text": text,
}
with open("page.json", "w", encoding="utf-8") as output:
    json.dump(record, output, ensure_ascii=False, indent=2)
print(f"Saved {len(text)} characters from {response.url}")

Install the two dependencies with python -m pip install requests beautifulsoup4. Replace the example host and URL with a page you are authorized to collect. The sample’s selector removal and short-page threshold are heuristics: inspect extracted results for your site before using them across a corpus.

9. Or skip the browser setup

When a page needs browser rendering and your goal is a visual capture, ScreenshotNeo returns an image or PDF from one API request. It accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. An MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Use a screenshot or PDF as a visual artifact; use extracted text and structured metadata for semantic RAG retrieval.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/docs/ -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/docs/"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/docs/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

For Node.js, the example uses the built-in fetch and Bun’s file writer; in Node.js, write the response body with await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))). The supported API options and parameter names are documented at ScreenshotNeo docs. Sign up for 1,000 free screenshots a month with no card.

10. Troubleshooting

Symptom Likely cause What to do
HTTP 401 or 403 Authentication is required, the crawler is not permitted, or the page has an access restriction. Confirm permission and supported authentication with the site owner. Do not try to bypass the restriction.
HTTP 429 or repeated throttling Request rate or concurrency is too high, or the site is limiting the crawler. Reduce concurrency, honor retry guidance, and use backoff. Revisit only changed pages where possible.
Successful response but no article text The page may require JavaScript, use unusual markup, or return a challenge/interstitial. Inspect the response and content type. Use an allowed rendered-browser workflow or supported API if appropriate; treat challenges as access boundaries.
Same article appears several times Redirect, trailing slash, query parameter, or hostname variants were treated as separate documents. Normalize URLs, validate canonical links, and deduplicate by stable page identity and content hash.
Passages lack context in answers Chunks split away headings, table headers, definitions, or conditions. Chunk on semantic boundaries and attach title and heading path; evaluate with real questions and inspect retrieved passages.
Old content keeps appearing Refresh did not replace prior chunks or deleted pages remain indexed. Use stable document IDs, version updates, and define deletion or unavailable-page behavior.
Fetch hangs or consumes too much memory No timeouts or size limits, or oversized pages are processed at once. Set connect/read timeouts and response-size limits; stream where suitable and isolate unusually large documents.
Canonical URL points outside intended scope The site declares a canonical target that is unrelated, cross-domain, or unsuitable for your corpus. Validate canonical targets against your host and content policy; retain the original URL for traceability.

11. Performance, reliability, and cost

For a first pass, the main cost drivers are pages fetched, browser rendering when needed, parsing, embedding tokens, index storage, and refresh frequency. No single price or crawl-rate estimate applies across sites or providers. Measure the target corpus and keep these costs visible by stage.

  • Reduce unnecessary work: deduplicate before fetching, use validators or hashes to skip unchanged pages, and avoid crawling irrelevant URL patterns.
  • Bound concurrency: higher parallelism can improve throughput but can also increase server load and throttling. Tune it to the site’s policy and observed responses.
  • Retry selectively: retry transient network failures and eligible server errors with bounded backoff; do not loop on authorization failures, invalid URLs, or access challenges.
  • Make ingestion resumable: persist queue state and per-page status so a worker restart does not restart the entire crawl.
  • Separate fetch from indexing: retain the last good document while a refresh is pending or fails, then atomically replace chunks after successful preparation.
  • Monitor quality as well as throughput: track fetch failures, empty extraction, duplicate rate, stale documents, and retrieval quality on a representative question set.

A hosted ingestion or parsing service may reduce infrastructure work, but compare options against your actual corpus: access behavior, URL normalization, refresh support, JavaScript and document formats, structure preservation, chunk quality, traceability, pacing, and maintenance. The cited sources describe capabilities and workflows, not a universal comparative benchmark. [Google Cloud data preparation; Google Cloud parsing and chunking]

12. FAQ

Does a sitemap mean I am allowed to scrape every listed URL?

No. A sitemap helps discover URLs. Check access instructions and permissions for the crawler and content.

Does blocking a path in robots.txt protect its contents?

No. Robots.txt communicates crawler preferences; it does not provide confidentiality or authenticate users. Use actual access controls for private content.

Should I embed the HTML source?

Usually, extract readable content and meaningful structure first. Raw HTML contains markup and repeated page chrome that can reduce retrieval quality.

How do I know whether chunking is good?

Test representative questions, inspect the retrieved source passages, and adjust boundaries or context when the correct page is missing or its passage is incomplete.

Can screenshots alone power semantic RAG?

Not reliably. Screenshots are visual records; semantic retrieval needs extracted text or a suitable OCR and document-processing step.

Sources