How to Build a Modular RAG MCP Server
Build an MCP server with replaceable ingestion, retrieval, storage, and generation modules, plus tools, transports, security, and operations guidance.
Short answer: Build MCP as a protocol adapter around ordinary RAG modules. Keep ingestion, chunking, embedding, storage, retrieval, reranking, and answer synthesis behind explicit interfaces. Expose read-only search and answer tools separately from document and collection mutation tools, then choose stdio or a network transport for the deployment.
MCP standardizes how an application discovers and calls server capabilities such as tools, resources, and prompts. It does not require a particular vector database, embedding model, LLM, or internal service topology. The official documentation describes MCP as separating the concern of providing context from the LLM interaction itself (MCP documentation).
1. Architecture: keep the protocol boundary thin
A maintainable server has two boundaries:
- MCP boundary: validates arguments, authorizes the caller, invokes an application operation, and maps the result to MCP output.
- RAG boundary: parses documents, creates chunks and embeddings, searches storage, optionally reranks results, and synthesizes an answer with source metadata.
Use either of these valid shapes:
| Shape | How it works | When it fits |
|---|---|---|
| Thin adapter | MCP handlers call separate RAG and ingestion APIs. | Independent deployment, scaling, or teams. NVIDIA’s RAG MCP guide documents this pattern. |
| Separated retrieval service | An agent handles reasoning while an MCP service handles knowledge-base construction and retrieval. Embedding, vector storage, and model services can be separate. | Large deployments or independently managed infrastructure. AMD’s Agentic RAG blueprint demonstrates this arrangement. |
| Single process | MCP handlers call local ingestion and retrieval classes. | Prototypes, local tools, and tests. The interfaces should still remain replaceable. |
These are architectural options, not MCP requirements. Keep contracts stable so you can replace the embedding model, vector store, or generator without rewriting tool handlers.
2. Define contracts before implementation
Pass structured records between modules. A useful minimum contract is:
Document { id, source, text, metadata }
Chunk { id, document_id, text, start_offset, metadata }
SearchResult { chunk_id, document_id, text, score, metadata }
Answer { text, sources[] }
Preserve a stable document identifier and location metadata through chunking and retrieval. That lets the answer layer produce citations that point back to the source instead of citing an anonymous vector match.
3. A complete Python MCP server
The example below uses an in-memory store and deterministic token-frequency scoring so it runs without a model account or database. Replace the InMemoryStore and SimpleRetriever implementations with a persistent vector store and embedding service when you move to production. The MCP Python SDK has version-sensitive APIs; its v1 maintenance documentation identifies v2 as the current stable line, so check the current Python SDK documentation before pinning a release.
Install and run
python -m venv .venv
. .venv/bin/activate
pip install "mcp[cli]"
python server.py
server.py
from __future__ import annotations
import re
from dataclasses import dataclass, field
from typing import Any
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("modular-rag")
def terms(value: str) -> set[str]:
return set(re.findall(r"[a-z0-9]{2,}", value.lower()))
@dataclass
class Chunk:
id: str
document_id: str
text: str
start_offset: int
metadata: dict[str, Any] = field(default_factory=dict)
class InMemoryStore:
def __init__(self) -> None:
self.documents: dict[str, dict[str, Any]] = {}
self.chunks: dict[str, Chunk] = {}
def add(self, document_id: str, text: str, source: str,
metadata: dict[str, Any] | None = None,
chunk_size: int = 900) -> int:
self.delete(document_id)
self.documents[document_id] = {
"id": document_id, "source": source,
"metadata": metadata or {}, "text_length": len(text)
}
words = text.split()
created = 0
offset = 0
for start in range(0, len(words), chunk_size):
part = " ".join(words[start:start + chunk_size])
chunk_id = f"{document_id}:{created}"
self.chunks[chunk_id] = Chunk(
id=chunk_id, document_id=document_id, text=part,
start_offset=offset,
metadata={"source": source, **(metadata or {})},
)
offset += len(part) + 1
created += 1
return created
def delete(self, document_id: str) -> bool:
existed = document_id in self.documents
self.documents.pop(document_id, None)
for chunk_id in [k for k, v in self.chunks.items()
if v.document_id == document_id]:
del self.chunks[chunk_id]
return existed
def search(self, query: str, limit: int = 5) -> list[dict[str, Any]]:
query_terms = terms(query)
scored: list[tuple[float, Chunk]] = []
for chunk in self.chunks.values():
chunk_terms = terms(chunk.text)
if not query_terms or not chunk_terms:
continue
overlap = len(query_terms & chunk_terms)
score = overlap / (len(query_terms) ** 0.5 * len(chunk_terms) ** 0.5)
if score > 0:
scored.append((score, chunk))
scored.sort(key=lambda item: item[0], reverse=True)
return [{
"chunk_id": chunk.id,
"document_id": chunk.document_id,
"text": chunk.text,
"score": round(score, 4),
"metadata": chunk.metadata,
"start_offset": chunk.start_offset,
} for score, chunk in scored[:limit]]
store = InMemoryStore()
@mcp.tool()
def search(query: str, limit: int = 5) -> dict[str, Any]:
"""Return source-aware chunks relevant to a query."""
if not query.strip():
raise ValueError("query must not be empty")
if not 1 <= limit <= 20:
raise ValueError("limit must be between 1 and 20")
return {"query": query, "results": store.search(query, limit)}
@mcp.tool()
def ask(query: str, limit: int = 5) -> dict[str, Any]:
"""Retrieve context and return an extractive answer with citations."""
results = store.search(query, limit)
if not results:
return {"answer": "I could not find supporting passages.", "sources": []}
answer = " ".join(result["text"] for result in results)
sources = [{
"document_id": result["document_id"],
"chunk_id": result["chunk_id"],
"source": result["metadata"].get("source"),
"score": result["score"],
} for result in results]
return {"answer": answer, "sources": sources}
@mcp.tool()
def add_document(document_id: str, source: str, text: str,
metadata: dict[str, Any] | None = None) -> dict[str, Any]:
"""Add or replace a document. Protect this tool with write authorization."""
if not document_id.strip() or not source.strip() or not text.strip():
raise ValueError("document_id, source, and text are required")
count = store.add(document_id, text, source, metadata)
return {"document_id": document_id, "chunks_created": count}
@mcp.tool()
def delete_document(document_id: str) -> dict[str, Any]:
"""Delete a document and all of its chunks. Protect this tool with write authorization."""
return {"document_id": document_id, "deleted": store.delete(document_id)}
@mcp.tool()
def stats() -> dict[str, int]:
"""Return basic collection counts."""
return {"documents": len(store.documents), "chunks": len(store.chunks)}
if __name__ == "__main__":
mcp.run()
The TypeScript SDK exposes a similar high-level McpServer abstraction for tools, resources, and prompts (TypeScript SDK). The important design decision is the contract between each handler and the RAG application, not the programming language.
4. Tool design: separate reads from writes
A minimal query surface has search for source-aware chunks and ask for retrieval plus synthesis. Add administrative tools only when clients need them:
create_collection,list_collections, andcollection_statsupload_document,update_document, anddelete_documentclear_collectionfor controlled maintenance
NVIDIA and AMD reference implementations expose combinations of search, generation, upload, update, delete, build, clear, and statistics operations. Treat those as examples. Keep destructive operations separate, clearly described, and restricted by identity and collection permissions.
5. Ingestion and retrieval data flow
- Read the source document and record its canonical identifier.
- Extract text while retaining page, heading, paragraph, or character offsets.
- Chunk text with overlap appropriate to the document structure.
- Create embeddings and store vectors with document and location metadata.
- Embed or otherwise process the query.
- Retrieve candidate chunks using similarity and metadata filters.
- Optionally deduplicate, rerank with a cross-encoder, or grade relevance.
- Give only selected context to the generator and return citations from metadata.
AMD’s blueprint describes iterative retrieval and relevance grading. A community implementation uses Chroma, optional Hugging Face cross-encoder reranking, and metadata-derived citations. These are implementation choices, not universal requirements.
6. Choosing storage, models, and service boundaries
| Decision | Option A | Option B | Trade-off |
|---|---|---|---|
| Deployment | Single process | Separate services | Simplicity versus independent scaling and release cycles. |
| Models | Local embedding and LLM services | Hosted OpenAI-compatible endpoints | Control and data locality versus operational convenience and network dependency. |
| Vector store | Embedded or local database | Managed or dedicated service | Lower setup cost versus durability, filtering, and operations managed elsewhere. |
| Retrieval | Vector similarity | Hybrid retrieval, reranking, or grading | Lower latency and complexity versus potentially better relevance at extra compute cost. |
| Tools | Query only | Query plus administration | Smaller attack surface versus self-service knowledge-base management. |
The cited architectures use different combinations, including ChromaDB, MMR retrieval, vLLM embeddings, OpenAI-compatible LLM endpoints, and cross-encoder reranking. No cited source establishes one database, transport, or model as best overall.
7. MCP transports and deployment
The Python SDK documentation lists stdio, SSE, and Streamable HTTP. NVIDIA documents all three, while AMD’s blueprint connects over SSE. Choose based on where the client runs:
- stdio: the client launches a local server process. It is convenient for desktop agents and local development.
- SSE: a network service that maintains server-sent event communication. Confirm current client support before deployment.
- Streamable HTTP: a network transport suitable for separately hosted services when supported by your selected SDK and client.
Pin an SDK version after checking its current documentation, and verify the chosen transport with the exact MCP client you will operate. Transport names and support can change between SDK releases.
8. Authentication and authorization
Authenticate at the network boundary and authorize inside the application. Validate tokens before dispatch, map identities to roles, and check collection permissions on every read and write. A MariaDB architecture example uses gateway token validation, role checks in the RAG API, and tool registration that adapts to service availability; treat it as an architectural pattern rather than a complete security standard.
- Give read-only identities access to
search,ask, and selected resources. - Require an ingestion role for uploads and updates.
- Require an administrator role for deletes, clears, and collection changes.
- Do not trust a document ID or collection name supplied by the model without authorization checks.
- Log caller, tool, collection, document ID, latency, result count, and failure reason without logging secrets.
9. Reliability, performance, and cost
Reliability checklist
- Make document upserts idempotent using a stable document ID and content hash.
- Persist ingestion status so interrupted jobs can resume.
- Set timeouts for embedding, vector-store, and generation calls.
- Return partial retrieval failures explicitly instead of presenting an unsupported answer.
- Keep source metadata with every chunk so citations survive reindexing.
- Use bounded result counts and context budgets to protect the generator.
Performance levers
- Batch embeddings during ingestion.
- Cache embeddings by content hash.
- Filter by tenant, collection, language, or document type before vector search.
- Retrieve a modest candidate set, then rerank only that set.
- Stream long generation responses when the client and transport support it.
- Measure ingestion time, retrieval latency, reranking latency, generation latency, token usage, and error rate separately.
The supplied research does not contain independent latency, accuracy, scaling, or cost benchmarks. Do not infer those numbers from architecture diagrams. Your cost model should include embedding calls, vector storage, model generation, network transfer, and the compute required for reranking.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Client cannot discover tools | Wrong transport, command, or SDK protocol version. | Run the server directly, inspect stderr, confirm the client’s transport configuration, and check the current SDK guide. |
| Tools appear but calls fail validation | Handler schema and client arguments differ. | Use typed parameters, clear descriptions, and reject invalid ranges with actionable errors. |
| Search returns nothing | Ingestion failed, query terms do not overlap, or metadata filters exclude every chunk. | Check collection stats, inspect one stored chunk, verify the embedding/index job, and test without filters. |
| Answers cite the wrong source | Chunk metadata was dropped or document IDs are unstable. | Persist source and location fields through chunking, retrieval, reranking, and synthesis. |
| Answers hallucinate | Generator received weak context or was allowed to answer without evidence. | Require supporting chunks, return an explicit no-answer result, and expose citations with scores and locations. |
| Writes are unauthorized | Authentication exists but authorization is only checked at the gateway. | Check role and collection permissions inside every mutating handler. |
| Latency grows sharply | Too many candidates, oversized context, synchronous ingestion, or remote model calls. | Bound limits, batch work, cache embeddings, rerank fewer candidates, and measure each stage separately. |
11. Test the modules independently
- Unit-test chunk boundaries, offsets, metadata preservation, and deterministic IDs.
- Test the retriever with known queries and expected source documents.
- Test authorization for every read, write, and destructive operation.
- Run ingestion and retrieval against the real storage adapter before adding MCP.
- Connect the exact MCP client and verify tool discovery, argument validation, errors, and transport shutdown.
- Test restart and retry behavior for interrupted ingestion and unavailable model services.
12. Or skip the browser setup
If your agent needs screenshots of documentation, dashboards, or rendered RAG results, ScreenshotNeo provides a single HTTP call instead of maintaining browser automation. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the other capture options. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
13. FAQ
Should the MCP server contain the entire RAG pipeline?
No. A thin adapter that forwards calls to ingestion and retrieval APIs is valid, as is a single-process implementation. Choose based on deployment and ownership boundaries.
Which transport should I use?
Use stdio for a client-launched local process. Use SSE or Streamable HTTP for a separately hosted service after confirming support in your current SDK and client.
Do I need a vector database?
Production systems usually need durable vector and metadata storage, but the protocol does not require one. The example uses an in-memory store so the interfaces are easy to understand.
What should an answer tool return?
Return the answer plus source identifiers and locations. A client can then display citations or let the user inspect the supporting documents.
How many tools should I expose?
Start with search and ask. Add ingestion and collection administration only when a client needs them, and protect mutation tools with separate roles.


