MCP Server Architecture
Understand MCP server architecture, transports, stateless lifecycle, tools, resources, prompts, security, scaling, and version migration.

MCP (Model Context Protocol) defines a standard communication surface between an AI application and servers that provide tools, resources, and prompts. The host application owns the user experience, an MCP client manages each server connection, and the MCP server maps protocol requests to your application logic and downstream systems.
The current baseline is MCP 2026-07-28. It has a stateless protocol core: requests carry their own protocol metadata, the legacy initialize/initialized handshake and Mcp-Session-Id are removed, and servers can scale behind ordinary load balancers. Older tutorials describe the 2025-11-25 lifecycle, so version every implementation and test mixed-version clients deliberately. Read the versioned specification before shipping.
What is an MCP server?
An MCP server is a process or network service that exposes capabilities through MCP. It does not have to contain your business logic; MCP standardizes discovery, schemas, messages, transport, and lifecycle behavior while your server calls databases, APIs, filesystems, browsers, or other services.
| Primitive | Purpose | Control model | Typical examples |
|---|---|---|---|
| Tools | Perform an operation | Model-controlled, subject to host policy and authorization | Query an order, create a ticket, capture a screenshot |
| Resources | Supply contextual data identified by a resource URI | Application-controlled | Files, records, documentation, schemas |
| Prompts | Reusable interaction templates | User-controlled | Review a pull request, summarize a report |
Keeping these control models separate matters. A prompt is not an executable permission. A resource is not an action. A tool should expose the smallest operation and data scope needed for the task.
How MCP server architecture fits together
- Host: the AI application (for example, an IDE or agent) owns conversation state, user consent, model calls, and policy.
- Client: an MCP client inside or alongside the host maintains a connection to one server and translates host decisions into protocol requests.
- Transport: stdio or Streamable HTTP carries JSON-RPC messages and binding metadata.
- Server: validates requests, checks authorization, invokes handlers, and returns bounded structured results.
- Downstream systems: APIs, databases, queues, browsers, or local files perform the actual work.
user -> AI application (host)
|
+-- MCP client ---- stdio ---- local MCP server ---- database/API
|
+-- MCP client ---- Streamable HTTP ---- remote MCP server ---- services
The host may connect to many servers. Each client-server relationship has its own capabilities and policy. Do not let a server assume that the model, host, or another server has authenticated a request; authenticate and authorize at the server boundary.

What changed in the 2026-07-28 lifecycle?
The 2026-07-28 release changes the architectural baseline described in many earlier examples:
- There is no protocol-level
initialize/initializedexchange in the current lifecycle. Mcp-Session-Idis removed. Each request is self-describing and can reach any healthy server instance.- An optional
server/discovercall lets a client inspect supported versions, capabilities, and server identity before normal requests. - When an application needs continuity, issue an explicit opaque handle as a tool result and require that handle in later arguments. A handle references protected state; it is not a credential.
- Streamable HTTP requires
Mcp-MethodandMcp-Nameheaders. Intermediaries can route or meter requests without parsing the JSON body, and servers must reject header/body mismatches according to the binding rules. - Legacy HTTP+SSE is deprecated with a minimum twelve-month deprecation window. Tasks are now an extension, and server-initiated input can use Multi Round-Trip Requests (MRTR).
Clients may discover the current version and fall back to a legacy initialize handshake for older servers. Record the versions you support, expose compatibility tests, and never treat a draft document as the specification for a released deployment.
Transport choices: stdio versus Streamable HTTP
| Decision | stdio | Streamable HTTP |
|---|---|---|
| Connection | Client launches a subprocess; newline-delimited JSON-RPC on stdin/stdout | HTTP POST to one MCP endpoint; response is JSON or a request-scoped SSE stream |
| Best fit | Local IDE integrations, private workstation tools | Remote services, gateways, shared infrastructure |
| Scaling | Scale by launching processes | Ordinary load balancing; no sticky protocol sessions required |
| Security boundary | Process normally has the client user’s privileges | Network authentication, authorization, TLS, and egress controls |
| Operational burden | Simple deployment, harder fleet observability | Standard HTTP monitoring, rate limits, proxies, and autoscaling |
Transport changes framing and metadata carriage, not protocol meaning. Do not confuse current Streamable HTTP with the older HTTP+SSE transport. Choose based on locality and data boundaries, latency, client compatibility, authentication, observability, and operating cost.

Minimal MCP server in TypeScript (stdio)
The official TypeScript SDK demonstrates a server registering a get-forecast tool with a structured input schema and serving it over stdio. The exact SDK API can change with the protocol version, so pin a release and consult its versioned documentation. The following shows the architecture and the JSON-RPC messages a client must be able to produce.
import { createInterface } from "node:readline";
// Minimal educational server: one JSON-RPC request per input line.
// In production, use the version-pinned official TypeScript SDK.
const rl = createInterface({ input: process.stdin, crlfDelay: Infinity });
function reply(id: unknown, result: unknown) {
process.stdout.write(JSON.stringify({ jsonrpc: "2.0", id, result }) + "\n");
}
rl.on("line", (line) => {
let request: any;
try { request = JSON.parse(line); }
catch { return; }
if (request.method === "tools/list") {
return reply(request.id, { tools: [{
name: "get_forecast",
description: "Return a weather forecast for a city.",
inputSchema: {
type: "object",
properties: { city: { type: "string", minLength: 1, maxLength: 100 } },
required: ["city"],
additionalProperties: false
}
}] });
}
if (request.method === "tools/call" && request.params?.name === "get_forecast") {
const city = request.params.arguments?.city;
if (typeof city !== "string" || city.length === 0 || city.length > 100) {
return reply(request.id, { isError: true, content: [{
type: "text", text: "city must be a non-empty string of 100 characters or fewer"
}] });
}
// Replace this deterministic placeholder with an authorized downstream call.
return reply(request.id, { content: [{
type: "text", text: JSON.stringify({ city, forecast: "sunny" })
}] });
}
reply(request.id, { isError: true, content: [{ type: "text", text: "Unknown method" }] });
});
Run a local server with the exact command visible to the user before the client launches it. Keep stdout reserved for protocol messages; send diagnostics to stderr. In production, validate JSON Schema arguments, enforce authorization inside each handler, cap output size, and use the SDK’s cancellation and error facilities.
Inspecting a Streamable HTTP server with cURL
HTTP binding details such as authentication and required headers depend on the server’s documented version. A conceptual request looks like this:
curl -i https://example.com/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-H 'Mcp-Method: tools/list' \
-H 'Mcp-Name: inventory' \
-H 'Authorization: Bearer YOUR_TOKEN' \
--data '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
Replace the endpoint, token, and server name with values from that server’s documentation. A production client must verify the response content type, parse either JSON or a request-scoped SSE stream, honor cancellation, and reject a response whose protocol metadata conflicts with its headers.
Calling an MCP-backed HTTP service from Python
import os
import requests
endpoint = os.environ["MCP_ENDPOINT"]
token = os.environ["MCP_TOKEN"]
payload = {
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list",
"params": {},
}
headers = {
"Accept": "application/json, text/event-stream",
"Content-Type": "application/json",
"Mcp-Method": "tools/list",
"Mcp-Name": "inventory",
"Authorization": f"Bearer {token}",
}
response = requests.post(endpoint, json=payload, headers=headers, timeout=30)
response.raise_for_status()
print(response.headers.get("content-type"))
print(response.text)
Calling an MCP-backed HTTP service from Node.js
const endpoint = process.env.MCP_ENDPOINT;
const token = process.env.MCP_TOKEN;
const body = {
jsonrpc: '2.0',
id: 1,
method: 'tools/list',
params: {}
};
const res = await fetch(endpoint, {
method: 'POST',
headers: {
'Accept': 'application/json, text/event-stream',
'Content-Type': 'application/json',
'Mcp-Method': 'tools/list',
'Mcp-Name': 'inventory',
'Authorization': `Bearer ${token}`
},
body: JSON.stringify(body)
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(res.headers.get('content-type'));
console.log(await res.text());
Designing tools, resources, and prompts
Tools
- Use stable, descriptive names and a precise description that tells a model when the operation is appropriate.
- Publish a strict input schema with required fields, bounds, enums, and
additionalProperties: falsewhere possible. - Validate again in the handler. Schema validation is not authorization.
- Return bounded, structured results. Paginate or summarize large datasets.
- Apply authorization to the operation and each object or tenant ID, then log the decision.
- Expose narrow task-specific tools instead of one large tool with many modes.
Resources
Give each resource a stable identifier and document freshness. The 2026-07-28 release adds ttlMs and cacheScope metadata to list/read responses so clients can make informed caching decisions. Use deterministic list ordering to keep catalogs and prompt caches stable.
Prompts
Prompts are reusable templates selected by a user or host. Keep them parameterized and safe to display before invocation. They should guide an interaction, not silently grant a tool permission.
State, tasks, and long-running work
Stateless protocol handling does not require a stateless application. If a workflow spans requests, store state in your normal database or queue and return an unpredictable handle. On every later request, authenticate the caller, look up the handle server-side, verify that it belongs to the caller and tenant, enforce expiry, and authorize the requested transition.
For work that exceeds a request timeout, use the Tasks extension’s task handles and polling operations. If a user answer is required, MRTR lets the server return an input_required result and lets the client retry with the answer. This avoids keeping a permanently open bidirectional stream. Keep task state separate from protocol identity and document retention and cancellation behavior.
Authentication and security architecture
Consider every MCP server a security boundary because its tools may reach sensitive APIs, data stores, or a local machine.
- No token passthrough: do not accept a token issued for another resource and forward it unchanged. Validate that credentials target your MCP server and enforce audience boundaries.
- OAuth confused deputy: identify the MCP client, request only needed downstream scopes, preserve per-client consent, validate redirect URIs exactly, and protect state and CSRF flows.
- Issuer validation: validate the authorization response issuer (
iss) under RFC 9207. Bind credentials to the issuer that minted them. Client ID Metadata Documents are the preferred direction; Dynamic Client Registration remains for compatibility but is deprecated. - SSRF: treat metadata URLs and redirects as untrusted. Require HTTPS in production, block private and reserved ranges where appropriate, validate redirect destinations, and apply egress controls.
- Handles: possession of an application handle is not authentication. Use high-entropy values, bind them to a verified principal, and expire them.
- Local execution: show the exact launch command and obtain consent before running an untrusted server. Use least privilege, sandboxing, and protected environment variables. Protect any local HTTP listener.
Scaling and reliability
- Keep protocol handlers request-scoped so any instance can serve any request.
- Put durable application state in a shared store only when the workflow needs it; do not recreate protocol sessions just to coordinate instances.
- Set deadlines for downstream calls, propagate cancellation, and return actionable error classes.
- Use idempotency keys for mutating tools that may be retried after a network failure.
- Rate-limit by authenticated principal and operation, and cap concurrency for expensive tools.
- Emit structured logs with request ID, server name, method, tool name, principal, latency, outcome, and downstream status. Never log bearer tokens or sensitive arguments.
- Test malformed JSON, unknown methods, schema violations, unauthorized object IDs, cancellation, duplicate requests, partial downstream failure, and mixed-version clients.
Migration checklist from older MCP tutorials
- Declare support for
2026-07-28and any legacy version you still accept. - Remove assumptions about
initialize,initialized, andMcp-Session-Idfrom the current path. - Add
server/discoverhandling where capability inspection improves client setup. - For Streamable HTTP, send and validate
Mcp-MethodandMcp-Name. - Replace legacy HTTP+SSE deployments or document the deprecation and an end date.
- Move long-running behavior to Tasks and interactive pauses to MRTR.
- Retest authorization, load balancing, retries, cancellation, and cache metadata.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Client waits for initialize forever | Client or server follows an older lifecycle | Pin the protocol version and use the 2026-07-28 discovery/request flow, or explicitly support the legacy fallback. |
| HTTP request rejected before JSON parsing | Missing or mismatched Mcp-Method/Mcp-Name |
Set headers from the actual method and server name; reject body/header mismatches consistently. |
| Tool appears but calls fail | Schema, authorization, or downstream validation error | Log a redacted request ID, validate arguments, check object-level permissions, and return a bounded actionable error. |
| Works locally, fails remotely | stdio assumptions, missing environment variables, TLS, proxy, or egress policy | Use Streamable HTTP for remote deployment, configure TLS and proxy limits, and test from the production network boundary. |
| State leaks between users | Opaque handle is treated as authentication or lookup is not tenant-bound | Authenticate every request, bind handles to the verified principal, use unpredictable values, and expire them. |
| Large responses exhaust the model context | Unbounded tool or resource output | Paginate, filter, summarize, cap bytes, and expose narrower tools. |
| Duplicate side effects after retry | Timeout occurred after downstream commit | Use idempotency keys and durable operation records. |
Performance, reliability, and cost notes
Protocol overhead is usually smaller than model and downstream-service latency, but tool design determines total cost. Keep schemas and descriptions concise, avoid sending unused resources, cache only when ttlMs and cacheScope permit it, and make list ordering deterministic. Measure server time, downstream time, serialization size, retries, and model tokens separately.
Stdio avoids a network hop and is often simplest for local tools. Streamable HTTP adds network and proxy behavior but enables shared deployment, horizontal scaling, standard TLS termination, and centralized observability. Stateless routing removes sticky-session infrastructure; it does not remove the need to pay for databases, queues, browser workers, or third-party APIs used by your tools.
Or skip the browser setup
If an MCP tool needs website screenshots, you can build and operate a browser worker yourself. ScreenshotNeo provides a website screenshot API and MCP server instead. One GET request returns PNG, JPEG, WebP, or PDF, and its MCP tools are take_screenshot, get_page_info, and capture_pdf.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. The MCP server lets Claude, Cursor, and other MCP clients take screenshots. Every plan includes the features; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000 shots. See the ScreenshotNeo API and MCP documentation, then sign up free.
FAQ
Is an MCP server the same as an AI agent?
No. The host runs the agent and model loop. An MCP server exposes bounded capabilities that the host may allow the model to call.
Do I need a database for a stateless MCP server?
No for purely request-scoped operations. Add normal application storage when workflows, handles, tasks, or audit records need continuity.
Can one host use multiple MCP servers?
Yes. The host can create one client connection per server and apply separate permissions, transports, and policies.
Should every capability be a tool?
No. Actions belong in tools, contextual data in resources, and reusable user-invoked templates in prompts.
Can I keep using HTTP+SSE?
It is deprecated in the 2026-07-28 release. Plan migration to Streamable HTTP and keep any compatibility path time-bounded.
Are protocol handles equivalent to API keys?
No. A handle references server-side state. Authenticate each request and bind the handle to the verified principal.


