ScreenshotNeo

BlogEngineering

Retrieval-Augmented Generation (RAG): Definition and How It Works

Learn what retrieval-augmented generation is, how retrieval and generation fit together, and how to build a reliable RAG workflow.

By the ScreenshotNeo team29 September 20268 min read

Retrieval-Augmented Generation (RAG): Definition and How It Works

Retrieval-augmented generation (RAG) is a system pattern that gives a language model access to an external collection of information at answer time. A retriever finds passages related to the user’s question, those passages are placed in the model’s context, and the model generates a response using both the retrieved evidence and its learned parameters. The external collection is separate from the model’s parametric memory, so it can be changed without retraining the entire model.

The pattern comes from the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In that work, a neural retriever searched a dense index of Wikipedia and supplied passages to a sequence-to-sequence generator. Modern RAG systems vary widely, but that division of labor—retrieve, then generate—is the defining idea.

RAG in one mental model

Think of a model’s parameters as learned memory and your corpus as a reference shelf. A normal prompt asks the model to answer from its learned memory. A RAG prompt first asks a search component to fetch likely relevant pages, chunks, or records from the shelf. The generator then sees the question plus the selected material.

RAG connects a question, retrieved evidence, and generation in sequence.
RAG connects a question, retrieved evidence, and generation in sequence.
  1. Corpus: documents, tickets, manuals, database rows, or another maintained collection are prepared for search.
  2. Question: the user’s input becomes a retrieval query, sometimes after rewriting.
  3. Retrieval: a sparse method such as TF-IDF/BM25, a dense dual-encoder, or hybrid search returns candidate passages. Dense vectors and vector databases are common choices, not requirements.
  4. Context assembly: the application filters, ranks, deduplicates, and formats the best passages with the question.
  5. Generation: the language model writes an answer conditioned on that context and its parameters.

Meta’s explainer describes the key change from a closed-book prompt: the input is used to retrieve relevant documents before generation. Retrieved text is evidence made available to the model, not proof that the final answer is correct.

How a RAG system works, step by step

1. Prepare and maintain the source

Start with a defined information boundary. Collect the files or records the application is allowed to use, preserve source identifiers and timestamps, remove duplicates, and extract text while retaining headings, tables, and links where they carry meaning. Split content into passages small enough to retrieve precisely but large enough to preserve context. Store metadata such as product version, language, permissions, and URL alongside every passage.

Indexing is separate from answering. When a source changes, reprocess affected passages and update the index. A maintained corpus can change without changing model weights, but it does not become current automatically: if the source is stale, retrieval is stale.

2. Turn a question into a search request

The raw question is often adequate. For multi-turn chat, resolve references such as “that limit” into a standalone query while preserving the original question for generation. Apply authorization filters before retrieval so a matching passage the user cannot access never enters the prompt.

3. Retrieve and rank candidates

Sparse retrieval matches terms and works well when exact names, error codes, or identifiers matter. Dense retrieval maps questions and passages into vectors and can match related wording. Dense Passage Retrieval reported a 9%–19% absolute improvement in top-20 passage retrieval accuracy over a strong Lucene-BM25 baseline on the datasets in that paper; treat that as an experiment-specific result, not a universal guarantee. Hybrid retrieval, metadata filters, and a second-stage reranker are common ways to balance recall and precision.

4. Build a bounded context

Do not blindly concatenate every hit. Select a small top-k, remove near-duplicates, enforce tenant and version filters, and fit the result within the model’s context window. Include source labels so the model and UI can cite where each statement came from. More retrieved text can raise token usage and inference cost under per-token billing.

5. Generate with explicit grounding instructions

Useful instructions tell the model to answer from supplied passages, distinguish missing evidence, and cite source identifiers. This reduces ambiguity but cannot force faithful use of context. The generator can misunderstand a passage, combine incompatible versions, or answer from prior knowledge when retrieval is weak.

A minimal, runnable Python RAG example

This example uses a tiny in-memory corpus, TF-IDF retrieval, and a placeholder generation function. Replace generate() with your model SDK.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

DOCS = [
    {'id': 'limits', 'text': 'The API allows 60 requests per minute per key.'},
    {'id': 'auth', 'text': 'Send the access key in the Authorization header.'},
    {'id': 'billing', 'text': 'A successful capture is billed after the image is returned.'},
]

vectorizer = TfidfVectorizer(stop_words='english')
X = vectorizer.fit_transform([d['text'] for d in DOCS])

def retrieve(question, k=2):
    q = vectorizer.transform([question])
    scores = cosine_similarity(q, X)[0]
    order = scores.argsort()[::-1][:k]
    return [DOCS[i] | {'score': float(scores[i])} for i in order]

def build_prompt(question, hits):
    evidence = '\n'.join(f"[{h['id']}] {h['text']}" for h in hits)
    return ('Answer only with support from EVIDENCE. If it is insufficient, say so.\n'
            f'EVIDENCE:\n{evidence}\nQUESTION: {question}')

def generate(prompt):
    return 'Pass prompt to your chosen language-model client.'

question = 'How many requests can one key make per minute?'
hits = retrieve(question)
print(build_prompt(question, hits))
print('Sources:', [h['id'] for h in hits])

For production, persist the index, add document-level permissions, record retrieval scores, and return citations with every answer. A low score or empty result should trigger a “not enough evidence” path instead of confident prose.

Design choices and configuration checklist

Decision Options What to check
Corpus Files, web pages, tickets, SQL rows Coverage, freshness, permissions, provenance
Chunking Fixed tokens, sentence/heading aware, parent-child Boundary quality and duplicate overlap
Retriever BM25/TF-IDF, dense, hybrid Recall on real queries and identifier handling
Ranking Similarity score, metadata filters, reranker Precision of the first few passages
Context Top-k, score threshold, token budget Conflicts, latency, and prompt cost
Generator Any model that accepts prompt and context Instruction following, citation format, refusal behavior
  • Keep an immutable source ID and revision on each chunk.
  • Log query, retrieved IDs, scores, model version, and latency.
  • Evaluate retrieval and generation separately.
  • Test empty, conflicting, multilingual, long, and adversarial questions.

What RAG improves—and what it cannot guarantee

RAG lets a generator consult information outside its parameters and gives an application a replaceable memory component. This can support private or domain-specific material and let maintainers update the corpus without retraining the whole model. The original paper reported state-of-the-art results on three open-domain question-answering tasks and more specific, diverse, and factual language than a parametric-only baseline in its evaluated generation tasks. Those are findings from that study, not promises for every deployment.

RAG does not guarantee factual answers, eliminate hallucinations, or make knowledge current. Failures come from missing documents, poor chunking, incorrect ranking, stale revisions, access-control bugs, or a generator that ignores or misreads context. Treat citations and abstention as product behavior, then verify it with representative evaluations.

Reliability, performance, and cost

Latency

Measure ingestion, retrieval, reranking, generation, and network time independently. Cache embeddings and repeated retrieval results where the corpus revision is unchanged. Parallelize independent searches and stream model output where appropriate.

Reliability

Use timeouts and retries with backoff, and make retries idempotent. Keep a degraded path: if retrieval is unavailable, either answer from a clearly labeled parametric mode or refuse, depending on risk. Store the exact passages used for an answer so incidents can be reproduced after the index changes.

Cost

Costs usually scale with indexed storage, retrieval infrastructure, reranking, and generated input/output tokens. Supplying more context can increase inference cost. Set a token budget, remove duplicate text, and compare quality at several k values rather than maximizing context.

Troubleshooting common RAG failures

Symptom Likely cause Fix
Covered content returns “I don’t know” Chunk missing, wrong filter, or query mismatch Inspect top-k IDs and scores; verify ingestion and filters; add rewriting or hybrid search.
Confident answer from the wrong document Low-precision retrieval or stale revision Add reranking, stricter thresholds, revision filters, and abstention.
Context exceeds model limit Top-k or chunk size too large Trim duplicates, lower k, summarize hierarchically, and reserve answer tokens.
Latency spikes Cold index, remote reranker, or oversized prompt Warm workers, cache frequent queries, parallelize retrieval, and cap context.
Users see restricted data Authorization applied after retrieval Apply tenant and ACL filters inside retrieval and test cross-tenant cases.
Updates are invisible Index not refreshed or cache keyed too broadly Track revisions, re-index changed records, and invalidate affected caches.

When to use RAG

RAG fits answers that depend on a changing or private corpus and need source-aware responses. It is less useful when no external evidence is needed, when a deterministic database query is the right interface, or when retrieval quality cannot meet the required risk threshold. Fine-tuning and RAG solve different problems; the research does not establish a universal rule that one replaces the other.

A capture pipeline can remove page overlays before storing visual evidence.
A capture pipeline can remove page overlays before storing visual evidence.

Or skip the browser setup

If your RAG project needs screenshots of documentation or rendered pages as corpus artifacts, ScreenshotNeo accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie/consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for full-page and element capture, custom CSS and JavaScript, waits, blocking rules, headers and cookies, device presets, PDF controls, caching, signed links, async webhooks, bulk capture, and usage reporting.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Does RAG always use a vector database?

No. Sparse indexes, dense indexes, hybrid systems, and ordinary database queries can all supply retrieval.

Is web search the same as RAG?

No. RAG can search a curated private corpus or application database. Web search is only one possible source.

Does adding more passages improve answers?

Not necessarily. Extra or conflicting text can distract the generator and raises context cost. Tune k and token budgets against evaluation queries.

How do I prove an answer is grounded?

Return source IDs and passages, log retrieval, and evaluate whether claims are entailed by those passages. A citation supports review but does not automatically prove correctness.