13 Best Web Data Sources for AI and LLM Training in 2026
Compare 13 web, code, scholarly and reference datasets for AI training, with selection criteria for quality, language, rights, freshness and cost.
The best web data source for AI or LLM training depends on your target language, domain, quality threshold, rights review and data-engineering capacity. There is no universal winner.
For maximum control, start with Common Crawl. For a processed English corpus, evaluate FineWeb. For multilingual coverage, examine FineWeb-2 and mC4. Add code, scholarly or reference sources when those domains matter to your model.
How to choose a training-data source
| Question | Why it matters |
|---|---|
| Raw or processed? | Raw archives provide control but require extraction, filtering, deduplication and auditing. Processed datasets encode someone else’s choices. |
| Which languages? | English-focused data is unsuitable as the only source for a multilingual model. Check language balance rather than relying on a dataset’s headline size. |
| Which domains? | Web text alone may underrepresent code, science, books or encyclopedic writing. |
| How recent? | A large corpus can still be based on older snapshots. Record crawl and processing dates. |
| What rights? | Review source-level terms, attribution, opt-out or removal policies, personal data documentation and commercial-use limits. |
| What can you operate? | Storage, transfer, parsing, deduplication and quality classification can cost more than the download itself. |
Quick recommendations
- Maximum breadth and control: Common Crawl.
- Large, filtered English web corpus: FineWeb.
- Educational English content: FineWeb-Edu, after checking the current release card and terms.
- Many languages: FineWeb-2 or an mC4 variant, after validating the live release.
- Mixed web, books, code and academic content: Dolma.
- Quality signals and deduplication identifiers: RedPajama-Data-V2.
- Code: The Stack v2, with repository-level license and removal review.
- Reference knowledge: Wikimedia project dumps.
- Scientific material: arXiv-related corpora such as S2ORC or peS2o, with publisher-rights review.
1. Common Crawl
Common Crawl is a broad archive and input source. Its data is stored through AWS Public Data Sets and academic cloud platforms. Choose it when you have the engineering capacity to select snapshots, extract useful text, remove boilerplate, deduplicate documents, filter unsafe or irrelevant material and preserve provenance.
Best for
Teams that need broad domain coverage, custom filters or reproducible snapshot selection.
Trade-offs
You must build much of the data pipeline yourself. Archive scale also creates substantial storage, transfer and processing requirements. Treat each crawl snapshot and source URL as provenance metadata.
2. FineWeb
FineWeb is an English Common Crawl-derived corpus with documented filtering and deduplication. Its current dataset card reports more than 18.5 trillion tokens from 96 dumps spanning summer 2013 through April 2024. The maintainers describe it as cleaned and deduplicated English web data and list ODC-By 1.0.
The April 2024 endpoint matters: the token count does not mean the corpus contains the newest web content. Review the card’s limitations and social-impact documentation before use.
Best for
English general-purpose pretraining when a processed corpus is preferable to raw Common Crawl.
3. FineWeb-Edu
FineWeb-Edu is an education-oriented FineWeb subset. It can be useful when instructional and educational material is central to the model’s purpose. Confirm the current card’s release size, filtering recipe, data composition and terms before committing to a training run.
Best for
Educational assistants, tutoring models and mixtures where educational quality is a deliberate objective.
4. FineWeb-2
FineWeb-2 extends the processing approach toward multilingual web data. Its 2025 paper reports a 20 TB, five-billion-document dataset covering more than 1,000 languages. Those figures describe the paper’s dataset; verify the live release, language balance and current access details before treating them as current.
Best for
Multilingual experiments that need a broad web-derived starting point.
5. C4 and mC4
C4 and mC4 are cleaned Common Crawl corpora. The available variants differ materially in filtering, language and how much content is retained. Read the relevant dataset card rather than assuming “C4” identifies one uniform collection.
Best for
Baseline experiments and multilingual mixtures where an established Common Crawl processing family is useful.
6. Dolma
Dolma combines web text with academic publications, code, books and encyclopedic material. AI2’s card reports a three-trillion-token dataset and describes v1.7 source-level statistics. The mixture can reduce dependence on web-only data, but each original source’s terms still apply.
Best for
General models that need several domains in one documented mixture.
7. RedPajama-Data-V2
RedPajama-Data-V2 documents 84 Common Crawl snapshots, more than 100 billion documents, quality signals for 30 billion documents and duplicate identifiers that can be used to form a 20-billion-document deduplicated collection. Its listed languages are English, German, French, Spanish and Italian.
Best for
Teams that want quality signals and duplicate metadata to construct their own filtered subset.
8. RefinedWeb
RefinedWeb is a Common Crawl-derived corpus associated with Falcon training. Compare its filtering pipeline, snapshot scope, hosting and access terms with FineWeb and other derivatives. The exact current hosted release should be verified against the upstream project before use.
9. DCLM-Baseline
DCLM-Baseline is a research-documented Common Crawl-derived baseline for general web pretraining. Treat it as a candidate to evaluate alongside FineWeb, C4, RefinedWeb, DolmaCC and RedPajama rather than as a proven universal winner. Confirm its current release card and terms.
10. The Stack v2
The Stack v2 is a code-focused source family. Code repositories carry heterogeneous licenses and source metadata, so perform repository-level license checks and follow opt-out or removal policies. A code corpus should normally complement, rather than silently replace, general text data.
Best for
Code generation, repair and software-engineering models that can preserve source provenance.
11. The Pile
The Pile is a mixed-source text corpus that can broaden a web-heavy mixture. Assess its constituent datasets, age and terms individually. The aggregate name does not settle the rights or quality of every component.
12. Wikimedia projects
Wikimedia dumps provide encyclopedic and reference content that can improve factual and knowledge-heavy coverage. Use the applicable project dump and follow its attribution and license requirements. Wikimedia is a focused complement, not a replacement for web-scale data.
13. arXiv and scholarly corpora
arXiv-related data and scholarly corpora such as S2ORC or peS2o add scientific and technical material. Confirm the corpus version, access terms and publisher rights for included papers. Academic text may require different deduplication and quality rules from ordinary web pages.
Build a defensible data mixture
- Write down the target tasks, languages, domains and evaluation sets.
- Select a raw source, a processed general corpus and any domain-specific supplements.
- Record snapshot dates, source URLs, dataset versions, hashes and processing code.
- Extract text while retaining document and origin metadata.
- Remove boilerplate, malware, spam, near duplicates and content outside the project’s policy.
- Run language identification and domain classification; inspect samples from every major bucket.
- Audit personal or sensitive data, licenses, attribution and removal requests.
- Deduplicate before tokenization and again across all component datasets.
- Measure mixture proportions and hold out evaluation sources to reduce leakage.
- Recheck licenses and source cards immediately before training and publication.
Minimal Common Crawl index workflow
The following Python example shows the shape of a reproducible fetch: obtain an index record, download a WARC response and then pass the payload to your own parser and policy filters. Index endpoints and collection names change, so use the current Common Crawl documentation when selecting a crawl.
import gzip
import io
import json
import requests
from warcio.archiveiterator import ArchiveIterator
index_url = "https://index.commoncrawl.org/CC-MAIN-2025-30-index"
params = {"url": "example.com/*", "output": "json", "filter": "status:200"}
record = requests.get(index_url, params=params, timeout=60).text.splitlines()[0]
meta = json.loads(record)
warc = requests.get(
"https://data.commoncrawl.org/" + meta["filename"],
headers={"Range": f"bytes={meta['offset']}-{int(meta['offset']) + int(meta['length']) - 1}"},
timeout=120,
)
for item in ArchiveIterator(io.BytesIO(warc.content), arc2warc=True):
if item.rec_type == "response":
html = item.content_stream().read()
print(len(html))
break
Pin the crawl identifier, save the index record and validate compressed-record handling in your own pipeline. Do not treat an HTTP 200 response as proof that the content is suitable for training.
Quality, rights and operations checklist
- Coverage: Is the language and domain distribution measured on your intended sample?
- Freshness: Are crawl and processing dates recent enough for the task?
- Duplicates: Are exact, near and cross-source duplicates removed?
- Contamination: Could benchmark or evaluation text be present?
- Safety: Are personal, sensitive, illegal or malicious materials handled by an explicit policy?
- Rights: Have component licenses, attribution and removal procedures been reviewed?
- Reproducibility: Can another engineer rebuild the exact shard from recorded versions and filters?
- Cost: Have object storage, egress, parsing, classification and reprocessing been estimated?
Validate source pages before you train
For a small audit sample, capture representative pages after your extraction rules run. Compare full-page and element-level captures, check dark-mode or locale variants, and verify that cookie banners, newsletter overlays and chat widgets are not being mistaken for content. Store the URL, timestamp, response status, parser result and capture verdict with each sample.
Or skip the browser setup
ScreenshotNeo provides a single GET request for PNG, JPEG, WebP or PDF captures. Cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. It also has an MCP server so Claude, Cursor and other MCP clients can call take_screenshot, get_page_info and capture_pdf.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the 63 capture options, including full-page lazy-image loading, CSS selectors, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture and PDF settings. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Performance, reliability and cost notes
- Prefer processed datasets when engineering time and repeated filtering would cost more than their convenience.
- Use streaming formats and shard sizes that let failed jobs restart without repeating an entire download.
- Cache immutable source objects and record checksums so preprocessing is repeatable.
- Sample early: a small, stratified audit can reveal boilerplate, language imbalance and licensing problems before a large run.
- Keep raw and filtered layers separate so a policy change does not require reacquiring the archive.
- For page captures, use caching and bulk or asynchronous jobs where appropriate; inspect verdict and billing headers rather than assuming every response is usable.
Common mistakes and fixes
| Problem | Likely cause | Fix |
|---|---|---|
| Corpus is huge but model quality is weak | Scale replaced filtering, deduplication or domain balance | Audit samples, improve quality signals and tune the mixture. |
| Training data is stale | Old crawl or processing cutoff | Record dates and add a newer snapshot or targeted source. |
| Multilingual performance is uneven | Language imbalance or noisy language identification | Measure tokens by language and rebalance verified subsets. |
| Legal review stalls late | Only the top-level dataset label was reviewed | Inspect component terms, provenance, attribution and removal policies before downloading at scale. |
| Pages are captured with overlays | Consent, newsletter or chat UI was treated as page content | Remove overlays in preprocessing or use a capture service that handles them before capture. |
| Capture returns a blank or blocked page | Bot check, timeout, failed load or site-specific behavior | Inspect the response verdict, retry with suitable waits or headers, and exclude the URL from training-quality samples. |
FAQ
Which dataset is best for training an LLM?
There is no single best dataset. Match the source to your language, domains, freshness, rights requirements and ability to process raw data.
Is Common Crawl ready to train on?
No. It is an archive and input source. Expect extraction, filtering, deduplication, safety review and provenance work.
Is FineWeb newer than every alternative?
No. Its current card reports large scale from crawls through April 2024. Always compare snapshot dates.
Does an open dataset grant unrestricted commercial rights?
No. Review component terms, attribution obligations, personal-data documentation and commercial-use limits.
Should I use only web data?
Usually not when code, science, books or reference knowledge are important to the model’s target tasks.
