Retrieval-Augmented Generation (RAG): Definition and How It Works
Learn what retrieval-augmented generation is, how retrieval and generation fit together, and how to build a reliable RAG workflow.

Retrieval-augmented generation (RAG) is a system pattern that gives a language model access to an external collection of information at answer time. A retriever finds passages related to the user’s question, those passages are placed in the model’s context, and the model generates a response using both the retrieved evidence and its learned parameters. The external collection is separate from the model’s parametric memory, so it can be changed without retraining the entire model.
The pattern comes from the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In that work, a neural retriever searched a dense index of Wikipedia and supplied passages to a sequence-to-sequence generator. Modern RAG systems vary widely, but that division of labor—retrieve, then generate—is the defining idea.
RAG in one mental model
Think of a model’s parameters as learned memory and your corpus as a reference shelf. A normal prompt asks the model to answer from its learned memory. A RAG prompt first asks a search component to fetch likely relevant pages, chunks, or records from the shelf. The generator then sees the question plus the selected material.

- Corpus: documents, tickets, manuals, database rows, or another maintained collection are prepared for search.
- Question: the user’s input becomes a retrieval query, sometimes after rewriting.
- Retrieval: a sparse method such as TF-IDF/BM25, a dense dual-encoder, or hybrid search returns candidate passages. Dense vectors and vector databases are common choices, not requirements.
- Context assembly: the application filters, ranks, deduplicates, and formats the best passages with the question.
- Generation: the language model writes an answer conditioned on that context and its parameters.
Meta’s explainer describes the key change from a closed-book prompt: the input is used to retrieve relevant documents before generation. Retrieved text is evidence made available to the model, not proof that the final answer is correct.
How a RAG system works, step by step
1. Prepare and maintain the source
Start with a defined information boundary. Collect the files or records the application is allowed to use, preserve source identifiers and timestamps, remove duplicates, and extract text while retaining headings, tables, and links where they carry meaning. Split content into passages small enough to retrieve precisely but large enough to preserve context. Store metadata such as product version, language, permissions, and URL alongside every passage.
Indexing is separate from answering. When a source changes, reprocess affected passages and update the index. A maintained corpus can change without changing model weights, but it does not become current automatically: if the source is stale, retrieval is stale.
2. Turn a question into a search request
The raw question is often adequate. For multi-turn chat, resolve references such as “that limit” into a standalone query while preserving the original question for generation. Apply authorization filters before retrieval so a matching passage the user cannot access never enters the prompt.
3. Retrieve and rank candidates
Sparse retrieval matches terms and works well when exact names, error codes, or identifiers matter. Dense retrieval maps questions and passages into vectors and can match related wording. Dense Passage Retrieval reported a 9%–19% absolute improvement in top-20 passage retrieval accuracy over a strong Lucene-BM25 baseline on the datasets in that paper; treat that as an experiment-specific result, not a universal guarantee. Hybrid retrieval, metadata filters, and a second-stage reranker are common ways to balance recall and precision.
4. Build a bounded context
Do not blindly concatenate every hit. Select a small top-k, remove near-duplicates, enforce tenant and version filters, and fit the result within the model’s context window. Include source labels so the model and UI can cite where each statement came from. More retrieved text can raise token usage and inference cost under per-token billing.
5. Generate with explicit grounding instructions
Useful instructions tell the model to answer from supplied passages, distinguish missing evidence, and cite source identifiers. This reduces ambiguity but cannot force faithful use of context. The generator can misunderstand a passage, combine incompatible versions, or answer from prior knowledge when retrieval is weak.
A minimal, runnable Python RAG example
This example uses a tiny in-memory corpus, TF-IDF retrieval, and a placeholder generation function. Replace generate() with your model SDK.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
DOCS = [
{'id': 'limits', 'text': 'The API allows 60 requests per minute per key.'},
{'id': 'auth', 'text': 'Send the access key in the Authorization header.'},
{'id': 'billing', 'text': 'A successful capture is billed after the image is returned.'},
]
vectorizer = TfidfVectorizer(stop_words='english')
X = vectorizer.fit_transform([d['text'] for d in DOCS])
def retrieve(question, k=2):
q = vectorizer.transform([question])
scores = cosine_similarity(q, X)[0]
order = scores.argsort()[::-1][:k]
return [DOCS[i] | {'score': float(scores[i])} for i in order]
def build_prompt(question, hits):
evidence = '\n'.join(f"[{h['id']}] {h['text']}" for h in hits)
return ('Answer only with support from EVIDENCE. If it is insufficient, say so.\n'
f'EVIDENCE:\n{evidence}\nQUESTION: {question}')
def generate(prompt):
return 'Pass prompt to your chosen language-model client.'
question = 'How many requests can one key make per minute?'
hits = retrieve(question)
print(build_prompt(question, hits))
print('Sources:', [h['id'] for h in hits])
For production, persist the index, add document-level permissions, record retrieval scores, and return citations with every answer. A low score or empty result should trigger a “not enough evidence” path instead of confident prose.
Design choices and configuration checklist
| Decision | Options | What to check |
|---|---|---|
| Corpus | Files, web pages, tickets, SQL rows | Coverage, freshness, permissions, provenance |
| Chunking | Fixed tokens, sentence/heading aware, parent-child | Boundary quality and duplicate overlap |
| Retriever | BM25/TF-IDF, dense, hybrid | Recall on real queries and identifier handling |
| Ranking | Similarity score, metadata filters, reranker | Precision of the first few passages |
| Context | Top-k, score threshold, token budget | Conflicts, latency, and prompt cost |
| Generator | Any model that accepts prompt and context | Instruction following, citation format, refusal behavior |
- Keep an immutable source ID and revision on each chunk.
- Log query, retrieved IDs, scores, model version, and latency.
- Evaluate retrieval and generation separately.
- Test empty, conflicting, multilingual, long, and adversarial questions.
What RAG improves—and what it cannot guarantee
RAG lets a generator consult information outside its parameters and gives an application a replaceable memory component. This can support private or domain-specific material and let maintainers update the corpus without retraining the whole model. The original paper reported state-of-the-art results on three open-domain question-answering tasks and more specific, diverse, and factual language than a parametric-only baseline in its evaluated generation tasks. Those are findings from that study, not promises for every deployment.
RAG does not guarantee factual answers, eliminate hallucinations, or make knowledge current. Failures come from missing documents, poor chunking, incorrect ranking, stale revisions, access-control bugs, or a generator that ignores or misreads context. Treat citations and abstention as product behavior, then verify it with representative evaluations.
Reliability, performance, and cost
Latency
Measure ingestion, retrieval, reranking, generation, and network time independently. Cache embeddings and repeated retrieval results where the corpus revision is unchanged. Parallelize independent searches and stream model output where appropriate.
Reliability
Use timeouts and retries with backoff, and make retries idempotent. Keep a degraded path: if retrieval is unavailable, either answer from a clearly labeled parametric mode or refuse, depending on risk. Store the exact passages used for an answer so incidents can be reproduced after the index changes.
Cost
Costs usually scale with indexed storage, retrieval infrastructure, reranking, and generated input/output tokens. Supplying more context can increase inference cost. Set a token budget, remove duplicate text, and compare quality at several k values rather than maximizing context.
Troubleshooting common RAG failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Covered content returns “I don’t know” | Chunk missing, wrong filter, or query mismatch | Inspect top-k IDs and scores; verify ingestion and filters; add rewriting or hybrid search. |
| Confident answer from the wrong document | Low-precision retrieval or stale revision | Add reranking, stricter thresholds, revision filters, and abstention. |
| Context exceeds model limit | Top-k or chunk size too large | Trim duplicates, lower k, summarize hierarchically, and reserve answer tokens. |
| Latency spikes | Cold index, remote reranker, or oversized prompt | Warm workers, cache frequent queries, parallelize retrieval, and cap context. |
| Users see restricted data | Authorization applied after retrieval | Apply tenant and ACL filters inside retrieval and test cross-tenant cases. |
| Updates are invisible | Index not refreshed or cache keyed too broadly | Track revisions, re-index changed records, and invalidate affected caches. |
When to use RAG
RAG fits answers that depend on a changing or private corpus and need source-aware responses. It is less useful when no external evidence is needed, when a deterministic database query is the right interface, or when retrieval quality cannot meet the required risk threshold. Fine-tuning and RAG solve different problems; the research does not establish a universal rule that one replaces the other.

Or skip the browser setup
If your RAG project needs screenshots of documentation or rendered pages as corpus artifacts, ScreenshotNeo accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie/consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for full-page and element capture, custom CSS and JavaScript, waits, blocking rules, headers and cookies, device presets, PDF controls, caching, signed links, async webhooks, bulk capture, and usage reporting.
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Does RAG always use a vector database?
No. Sparse indexes, dense indexes, hybrid systems, and ordinary database queries can all supply retrieval.
Is web search the same as RAG?
No. RAG can search a curated private corpus or application database. Web search is only one possible source.
Does adding more passages improve answers?
Not necessarily. Extra or conflicting text can distract the generator and raises context cost. Tune k and token budgets against evaluation queries.
How do I prove an answer is grounded?
Return source IDs and passages, log retrieval, and evaluate whether claims are entailed by those passages. A citation supports review but does not automatically prove correctness.


