ScreenshotNeo

BlogEngineering

Building an LLM-Ready Stack Exchange Corpus with a Crawling API

Plan a compliant Stack Exchange corpus: permissions, API and dump routes, attribution, licensing, ingestion code, and operational safeguards.

By the ScreenshotNeo team1 October 20268 min read

Short answer: Do not begin by crawling Stack Exchange. First write down the intended use, verify that the use is authorized, choose an approved access route, and design attribution and license handling before collecting or storing records. Stack Exchange’s Acceptable Use Policy prohibits automated data gathering from its network for developing, building, training, testing, indexing, benchmarking, or improving generative AI, chatbots, large language models, machine-learning systems, or similar systems unless you have express prior written consent. An API key or publicly visible page does not grant that permission.

1. Start with a permission decision tree

  1. Define the purpose. Is this research, search, evaluation, model training, a commercial product, or redistribution? Record whether the output will be public, private, or sold.
  2. Check the current rules. Read the Acceptable Use Policy, API Terms of Use, and Public Network Terms.
  3. If automated generative-AI collection is involved, obtain written permission first. The policy exception requires express prior written consent. Pause collection until that consent is documented.
  4. Choose the route that matches the approved use. Use the API for incremental, selective retrieval; use an official data dump when its terms and commercial status fit; use Data Explorer only after checking its current operational and reuse rules. Do not treat direct website crawling as the default ingestion method.
  5. Decide redistribution boundaries. A private corpus, embeddings, evaluation set, hosted search index, and redistributed dataset can have different obligations. Obtain qualified legal review for the exact derivative you plan to ship.

The API is programmatic access under terms. It is not blanket permission for every downstream use.

2. Compare the available access routes

Route Useful when What to verify
Stack Exchange API v2.3 You need selected fields, incremental windows, or a smaller corpus. API terms, attribution display, request keys or OAuth, filters, throttles, and conservative polling. The API returns JSON and supports field filters; semantically identical polling faster than once per minute is considered abusive. See the API documentation.
Creative Commons Data Dump You need a periodic snapshot and can comply with its license and access terms. The official staff announcement describes a new dump every three months and free access for non-commercial use, with commercial users directed to contact Stack Overflow. Confirm the current announcement and commercial agreement before downloading.
Data Explorer (SEDE) You need query-based exploration or a narrow export. Current export limits, update schedule, site coverage, and reuse terms. These operational details must be verified before depending on them.
Direct website crawling Only when you have explicit written permission that covers the proposed activity. Policy scope, request volume, load, retention, and redistribution. The Acceptable Use Policy bars automated extraction for generative-AI development without express prior written consent.

3. Design a record that preserves provenance

Keep source and license metadata with every record, including after transformations. A practical schema is an implementation recommendation, not a claim that Stack Exchange mandates these exact columns.

{
  'source_site': 'stackoverflow',
  'post_id': 123456,
  'post_type': 'question',
  'url': 'https://stackoverflow.com/questions/123456',
  'title': '...',
  'body_html': '...',
  'body_text': '...',
  'author_display_name': '...',
  'author_profile_url': '...',
  'content_license': 'CC BY-SA (verify version and terms)',
  'retrieved_at': '2026-10-01T12:00:00Z',
  'api_version': '2.3',
  'source_response_hash': '...',
  'transformations': ['html_to_text', 'code_blocks_preserved']
}

Keep the original URL, author attribution, source site, retrieval time, and applicable license when producing chunks, embeddings, evaluations, or exports. Store a transformation log so a reviewer can trace a derived record back to the source response.

4. Retrieve questions through the API

The example below fetches recent Stack Overflow questions in pages, respects the API’s backoff value, and stops when has_more is false. Replace the selection and filtering logic with the scope approved for your project.

Python

import json
import time
from datetime import datetime, timezone
from pathlib import Path

import requests

API = 'https://api.stackexchange.com/2.3/questions'
OUT = Path('questions.jsonl')

session = requests.Session()
session.headers.update({'User-Agent': 'approved-corpus-ingest/1.0'})

page = 1
while True:
    response = session.get(
        API,
        params={
            'site': 'stackoverflow',
            'pagesize': 100,
            'page': page,
            'sort': 'creation',
            'order': 'desc',
            'filter': 'default'
        },
        timeout=30
    )
    response.raise_for_status()
    payload = response.json()

    with OUT.open('a', encoding='utf-8') as handle:
        for item in payload.get('items', []):
            record = {
                'source_site': 'stackoverflow',
                'retrieved_at': datetime.now(timezone.utc).isoformat(),
                'api_version': '2.3',
                'post_id': item.get('question_id'),
                'url': item.get('link'),
                'title': item.get('title'),
                'author_display_name': (item.get('owner') or {}).get('display_name'),
                'raw': item
            }
            handle.write(json.dumps(record, ensure_ascii=False) + '\\n')

    if 'backoff' in payload:
        time.sleep(int(payload['backoff']))
    if not payload.get('has_more'):
        break
    page += 1
    time.sleep(1)  # Keep polling conservative; do not issue identical requests faster than once per minute.

cURL

curl --fail --silent --show-error \
  --get 'https://api.stackexchange.com/2.3/questions' \
  --data-urlencode 'site=stackoverflow' \
  --data-urlencode 'pagesize=100' \
  --data-urlencode 'page=1' \
  --data-urlencode 'sort=creation' \
  --data-urlencode 'order=desc' \
  --data-urlencode 'filter=default'

Node.js

const params = new URLSearchParams({
  site: 'stackoverflow',
  pagesize: '100',
  page: '1',
  sort: 'creation',
  order: 'desc',
  filter: 'default'
});

const response = await fetch(`https://api.stackexchange.com/2.3/questions?${params}`);
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const payload = await response.json();
for (const item of payload.items ?? []) {
  console.log(JSON.stringify({
    post_id: item.question_id,
    url: item.link,
    title: item.title,
    author_display_name: item.owner?.display_name,
    retrieved_at: new Date().toISOString()
  }));
}
if (payload.backoff) await new Promise(resolve => setTimeout(resolve, payload.backoff * 1000));

5. Select fields and collection boundaries deliberately

  • Use the narrowest documented endpoint and filter that satisfies the approved purpose.
  • Partition by site, date window, tags, or post type so a failed run can resume without duplicating the corpus.
  • Persist the request parameters and response metadata with each batch.
  • Respect API backoff instructions and avoid repeated polling. Cache responses locally when your terms and retention policy allow it.
  • Do not assume that collecting only titles, public profiles, or accepted answers removes policy or license obligations.

6. Attribution, licensing, and redistribution

API applications must visually identify Stack Exchange as the content source, and API use is subject to the API Terms of Use and Public Network Terms. The Public Network Terms identify the Creative Commons Data Dump as CC BY-SA; evaluate attribution and share-alike implications for your intended reuse. Keep attribution in the corpus, generated documents, evaluation reports, and any user-facing search or answer product where applicable.

Before distributing a dump, embeddings, or a trained artifact, ask:

  • Does the planned distribution include original text, excerpts, or reconstructable content?
  • What attribution travels with each record and each generated view?
  • Does the applicable CC BY-SA version impose share-alike duties on this derivative?
  • Does commercial use require a separate agreement?
  • Can a takedown, correction, or license change be mapped to stored records and derivatives?

7. Freshness, reliability, and cost planning

Freshness

API ingestion can be incremental but depends on your schedule and request budget. The dump is periodic; the official staff announcement describes a three-month cadence. Neither statement gives a guarantee that every site, field, or deleted item has the coverage your project needs.

Reliability

  • Checkpoint after every page or batch.
  • Use idempotent keys such as (source_site, post_id, revision).
  • Retry transient network failures with capped exponential backoff, while honoring API-provided backoff values.
  • Record HTTP status, quota metadata, request parameters, and a content hash.
  • Run validation for malformed JSON, missing IDs, duplicate records, and unexpected schema changes.

Cost

The dossier does not establish current numerical API quotas, dump sizes, or Data Explorer export limits. Budget storage, parsing, embedding, and review separately, and verify current limits before committing to a volume. Do not use an assumed quota or corpus-size figure in a capacity plan.

8. Troubleshooting

Symptom Likely cause Fix
Requests return an error or throttling response. Too many requests, repeated polling, or ignored backoff. Read the response metadata, honor backoff, slow the schedule, and cache completed pages.
The corpus contains duplicate posts. Jobs restarted without checkpoints or overlapping date windows. Upsert by site and post ID, retain batch manifests, and make retries idempotent.
Fields are missing. The selected filter does not expose them, or the endpoint does not return them. Check the v2.3 documentation and request only documented fields needed for the approved use.
A commercial launch is blocked. The chosen dump terms describe non-commercial access, or the generative-AI purpose lacks written consent. Stop collection and contact Stack Overflow for the applicable commercial agreement or written permission.
Attribution disappears after chunking. Metadata was stored only in a side table. Include source URL, author, license, and retrieval metadata in every exported record or maintain a verifiable join key.
A model or embedding index cannot support a takedown. No lineage from derivative to source record. Store hashes, record IDs, transformation versions, and deletion workflows before indexing.

9. Or skip the browser setup

If your workflow also needs visual snapshots of source pages or documentation, ScreenshotNeo provides a single screenshot API request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stackoverflow.com/questions -o shot.webp

Python

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stackoverflow.com/questions'}, timeout=90)
open('shot.webp', 'wb').write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stackoverflow.com/questions' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

10. FAQ

Can I crawl Stack Exchange for an LLM dataset?

Not by default. Automated gathering for generative-AI development requires express prior written consent under the Acceptable Use Policy.

Can I use the Stack Exchange API to build a corpus?

Use it only for an authorized purpose and comply with API terms, attribution requirements, request limits, and any written permission required for that purpose.

Is the Stack Exchange data dump free for commercial use?

The staff announcement describes free non-commercial access and directs commercial users to contact Stack Overflow. Verify current commercial terms before use.

How do I attribute Stack Exchange posts in a dataset?

Preserve the source URL, author attribution, source site, applicable license, and retrieval metadata at record level, then carry that information into derivatives and user-facing outputs as required.

Are embeddings automatically permitted?

No conclusion follows from technical transformation alone. Review the applicable policy, API terms, license, and intended distribution with qualified counsel.