ScreenshotNeo

BlogAI agents

How to Build an MCP Router in Python

Build a Python MCP router that aggregates tools from multiple servers, handles transports, names, failures, security, and deployment.

By the ScreenshotNeo team1 October 202611 min read

Direct answer: An MCP router is an MCP server to its caller and an MCP client to each downstream server. In Python, create one asynchronous client per backend, discover its tools, publish namespaced tool definitions, route calls through a lookup table, and forward results and errors without hiding failures.

The official Model Context Protocol specification defines hosts, clients, servers, and JSON-RPC messages. The current Python SDK documentation uses the v2 release line and requires Python 3.10 or newer. SDK package versions and negotiated MCP protocol versions are separate: installing SDK v2 does not force every peer to speak the newest protocol revision.

1. Decide what your router will expose

Start by writing down the router’s contract:

  • Tools: model-selected actions such as reading a file or running a search.
  • Resources: read-only data selected by the application.
  • Prompts: named prompt templates.
  • Backends: local subprocesses over stdio, deployed services over Streamable HTTP, or compatibility-only SSE services.
  • Failure policy: whether one unavailable backend removes only its tools or prevents startup.

This guide implements a tool router. Add resources and prompts only after defining how their URIs, names, subscriptions, and authorization rules map across backends.

2. Create the Python project

mkdir mcp-router
cd mcp-router
python3.10 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "mcp[cli]>=2,<3"

The mcp[cli] extra supplies development commands. Install the plain SDK package when you do not need those tools. Pin the major version in your application and review the SDK changelog before upgrades.

3. Understand the router data flow

  1. The upstream host connects to your router as an MCP client.
  2. The router opens one SDK Client for every configured backend.
  3. It calls each backend’s tool-list operation and builds a catalog.
  4. Every public name receives a stable backend prefix such as files__read_file.
  5. When the host calls a public name, the router resolves the mapping and invokes the original backend tool.
  6. The router returns the downstream content, structured result, and error state.

Namespacing is an implementation choice, not a protocol requirement. It prevents collisions and makes the tool’s origin visible to the model and to operators.

4. Implement the routing core

The following module is transport-neutral. It is fully runnable with an in-memory backend and provides the production behavior you need around discovery, namespacing, stale catalogs, and faithful error forwarding.

from __future__ import annotations

import asyncio
from dataclasses import dataclass
from typing import Any, Protocol


@dataclass(frozen=True)
class ToolSpec:
    name: str
    description: str
    input_schema: dict[str, Any]


@dataclass
class ToolResult:
    content: list[dict[str, Any]]
    structured_content: dict[str, Any] | None = None
    is_error: bool = False


class Backend(Protocol):
    async def list_tools(self) -> list[ToolSpec]: ...
    async def call_tool(self, name: str, arguments: dict[str, Any]) -> ToolResult: ...
    async def close(self) -> None: ...


@dataclass
class Binding:
    backend_id: str
    original_name: str
    spec: ToolSpec


class Router:
    def __init__(self, backends: dict[str, Backend]):
        self.backends = backends
        self.bindings: dict[str, Binding] = {}
        self.backend_errors: dict[str, str] = {}

    async def refresh_catalog(self) -> list[ToolSpec]:
        new_bindings: dict[str, Binding] = {}
        errors: dict[str, str] = {}

        for backend_id, backend in self.backends.items():
            try:
                tools = await backend.list_tools()
            except Exception as exc:
                errors[backend_id] = f"{type(exc).__name__}: {exc}"
                continue

            for spec in tools:
                public_name = f"{backend_id}__{spec.name}"
                if public_name in new_bindings:
                    raise RuntimeError(f"duplicate public tool name: {public_name}")
                new_bindings[public_name] = Binding(
                    backend_id=backend_id,
                    original_name=spec.name,
                    spec=ToolSpec(
                        name=public_name,
                        description=f"[{backend_id}] {spec.description}",
                        input_schema=spec.input_schema,
                    ),
                )

        # Replace the catalog atomically after all successful discovery calls.
        self.bindings = new_bindings
        self.backend_errors = errors
        return [binding.spec for binding in new_bindings.values()]

    async def call(self, public_name: str, arguments: dict[str, Any]) -> ToolResult:
        binding = self.bindings.get(public_name)
        if binding is None:
            return ToolResult(
                content=[{"type": "text", "text": f"Unknown tool: {public_name}"}],
                is_error=True,
            )

        backend = self.backends.get(binding.backend_id)
        if backend is None:
            return ToolResult(
                content=[{"type": "text", "text": f"Backend unavailable: {binding.backend_id}"}],
                is_error=True,
            )

        try:
            # Preserve the backend's content, structured result, and error flag.
            return await backend.call_tool(binding.original_name, arguments)
        except Exception as exc:
            return ToolResult(
                content=[{"type": "text", "text": f"Backend call failed: {type(exc).__name__}: {exc}"}],
                is_error=True,
            )

    async def close(self) -> None:
        await asyncio.gather(*(backend.close() for backend in self.backends.values()))


class DemoBackend:
    async def list_tools(self) -> list[ToolSpec]:
        return [ToolSpec(
            name="echo",
            description="Return the supplied message",
            input_schema={
                "type": "object",
                "properties": {"message": {"type": "string"}},
                "required": ["message"],
            },
        )]

    async def call_tool(self, name: str, arguments: dict[str, Any]) -> ToolResult:
        if name != "echo":
            return ToolResult(content=[{"type": "text", "text": "Unknown backend tool"}], is_error=True)
        return ToolResult(content=[{"type": "text", "text": str(arguments["message"])}])

    async def close(self) -> None:
        return None


async def main() -> None:
    router = Router({"demo": DemoBackend()})
    tools = await router.refresh_catalog()
    print([tool.name for tool in tools])
    result = await router.call("demo__echo", {"message": "hello"})
    print(result.content, result.is_error)
    await router.close()


if __name__ == "__main__":
    asyncio.run(main())

Run it with python router_core.py. In production, replace DemoBackend with adapters around the SDK’s asynchronous Client.

5. Connect downstream MCP servers

Streamable HTTP

Use Streamable HTTP for deployed backends. The SDK client supports HTTP headers, authentication, proxies, timeouts, and connection limits through its HTTP transport stack. Configure the exact endpoint URL where possible. Redirects across origins are rejected, and HTTPS-to-HTTP downgrade redirects are not followed.

from mcp import Client


async def http_backend(url: str, headers: dict[str, str] | None = None):
    # Use the SDK's Streamable HTTP client transport for your pinned v2 release.
    # Keep the connection inside an async context and pass explicit auth headers.
    client = Client(url=url, headers=headers or {})
    await client.connect()
    return client

Exact constructor and transport helper names can change between SDK releases. Check the v2 client reference linked from the official repository when pinning a version.

stdio subprocesses

Use stdio when the host launches a local server process. MCP JSON-RPC uses stdin and stdout, so stdout must contain protocol messages only. Send logs to stderr. Pass required credentials explicitly; child processes receive a minimal environment allow-list rather than an assumption that the complete parent environment is inherited.

from mcp import Client
from mcp.client.stdio import StdioServerParameters


async def stdio_backend(command: str, args: list[str], api_key: str):
    params = StdioServerParameters(
        command=command,
        args=args,
        env={"DOWNSTREAM_API_KEY": api_key},
    )
    client = Client(server_params=params)
    await client.connect()
    return client

SSE compatibility

SSE remains available for servers that have not migrated, but the SDK documentation describes it as superseded by Streamable HTTP in the 2025-03-26 protocol revision. Do not choose SSE for a new deployment unless compatibility requires it.

6. Adapt an SDK client to the routing core

Keep the SDK-specific code in a small adapter. The adapter should translate typed SDK results into the router’s internal ToolSpec and ToolResult objects, while preserving structured content and the SDK’s error flag.

class McpClientBackend:
    def __init__(self, client):
        self.client = client

    async def list_tools(self) -> list[ToolSpec]:
        response = await self.client.list_tools()
        return [
            ToolSpec(
                name=tool.name,
                description=tool.description or "",
                input_schema=tool.inputSchema,
            )
            for tool in response.tools
        ]

    async def call_tool(self, name: str, arguments: dict[str, Any]) -> ToolResult:
        response = await self.client.call_tool(name, arguments)
        return ToolResult(
            content=[item.model_dump() if hasattr(item, "model_dump") else item for item in response.content],
            structured_content=getattr(response, "structuredContent", None),
            is_error=bool(getattr(response, "isError", False)),
        )

    async def close(self) -> None:
        await self.client.close()

Field spellings can differ between SDK model versions. Confirm the response model in the version you pin. Always check the error flag before trusting structured content.

7. Publish the router as an MCP server

The public side uses the SDK’s server API. Register a typed function for each public tool or implement a dynamic dispatch tool if your server layer supports runtime catalogs. A typed function and its docstring allow the SDK to derive an input schema from type hints.

from mcp.server import MCPServer

server = MCPServer("python-router")
router = Router({})


@server.tool()
async def call_router_tool(name: str, arguments: dict[str, Any]) -> dict[str, Any]:
    """Call a namespaced tool exposed by a configured downstream MCP server."""
    result = await router.call(name, arguments)
    return {
        "content": result.content,
        "structured_content": result.structured_content,
        "is_error": result.is_error,
    }

For a production server, wire the SDK’s selected stdio or Streamable HTTP runner around server, refresh the catalog before accepting calls, and expose a health method that reports unavailable backends and the catalog timestamp. The SDK’s HTTP server is a protocol implementation, not a complete application server; use an ASGI server and process manager for deployment.

8. Choose a catalog strategy

Strategy Advantages Trade-offs
Refresh at startup Simple and predictable New backend tools require a restart
Periodic refresh Changes appear automatically Calls can race with catalog updates
Refresh on demand Useful for infrequent changes First call may be slower
Static configuration Stable schemas and low overhead Drift must be managed manually

Use atomic replacement: build a complete new mapping, then swap it in one operation. Keep the previous catalog if a refresh fails, and expose the refresh error to operators.

9. Handle failures and retries

  • Return an MCP error when a public name is unknown.
  • Return an error when the backend is unavailable instead of reporting success.
  • Apply timeouts per backend call so one server cannot block the router indefinitely.
  • Retry only operations you know are safe to repeat. Tool calls may have side effects.
  • Use bounded exponential backoff with jitter for connection recovery.
  • Limit concurrent calls per backend to protect both the router and downstream service.
  • Record backend ID, public tool name, duration, and outcome without logging secrets or arguments that contain sensitive data.

The MCP SDK does not prescribe a universal retry policy, cache lifetime, or failure-isolation design. Make those choices explicit in your application documentation.

10. Secure the router

Treat downstream metadata and tool descriptions as untrusted unless the server is trusted. Preserve user consent and authorization boundaries instead of silently using broad router credentials on behalf of every caller.

  • Allow-list backend endpoints and subprocess commands.
  • Store credentials in a secret manager or environment injected by the process supervisor.
  • Use separate credentials per backend where possible.
  • Validate tool arguments at the boundary and enforce maximum sizes.
  • Redact authorization headers, cookies, tokens, and private arguments from logs.
  • Configure allowed hosts and origins for real deployment hostnames.
  • Configure proxy headers correctly when TLS terminates before the application.
  • Remember that the SDK’s in-process subscription bus does not synchronize notifications across replicas; use an external implementation if multi-replica subscriptions are required.

11. Performance and cost considerations

  • Reuse long-lived downstream client connections instead of reconnecting for every call.
  • Discover tools concurrently, but cap concurrency when many backends are configured.
  • Cache the catalog for a defined interval and refresh it in the background when backend schemas are stable.
  • Keep public descriptions concise because every description consumes host context.
  • Namespace names deterministically so caches and logs remain useful across restarts.
  • Measure discovery latency, call latency, timeout rate, error rate, and backend connection count.
  • Run multiple workers only when the router state and notification model support it; in-memory catalogs are local to each worker.

Your infrastructure cost is driven by process count, network traffic, downstream services, and any model or tool provider charges. The MCP protocol itself does not define a pricing model.

12. Troubleshooting

Symptom Likely cause Fix
Nothing appears on stdout except logs stdio protocol output is mixed with logging Send logs to stderr and reserve stdout for JSON-RPC.
Backend tools disappear after refresh Discovery replaced the catalog after a partial failure Build a new catalog separately, keep the last known-good mapping, and report backend errors.
Two tools have the same public name No collision strategy or unstable prefix Use a stable backend prefix and fail discovery on duplicate names.
Calls hang indefinitely No per-call or transport timeout Set bounded connection and call timeouts; cancel tasks on deadline.
Structured output is trusted despite failure The caller ignored the result error flag Check isError before processing structured content.
HTTP connection fails after redirect Cross-origin or HTTPS downgrade redirect Configure the final endpoint URL directly and keep HTTPS.
Child process cannot authenticate Credential was assumed to be inherited Pass the required environment variables explicitly in stdio parameters.
Only one replica sees updates Catalog or subscription state is in process memory Use shared storage or an external notification implementation.
Import errors after upgrading v1 and v2 APIs were mixed Pin one major SDK line and update imports against its official documentation.

13. Test the router before deployment

  • Register two backends with identical tool names and verify namespacing.
  • Return a backend error and verify the upstream response remains an error.
  • Stop one backend during discovery and during a call.
  • Send malformed and oversized arguments.
  • Force a timeout and confirm cancellation does not leak a task.
  • Restart a backend and verify catalog refresh behavior.
  • Run the router through the same stdio or Streamable HTTP transport used by the host.
  • Verify that secrets never appear in logs or tool descriptions.

Or skip the browser setup

If your router or agent needs website screenshots, ScreenshotNeo provides a single HTTP endpoint instead of requiring you to operate a browser backend. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await Bun.write('shot.webp', buffer);

ScreenshotNeo supports full-page capture, CSS element capture, dark mode, device presets, custom viewports, retina scale, PDFs, custom CSS and JavaScript, clicks, waits, blocked resource types, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs, bulk capture, and a usage API. You can start with 1,000 screenshots per month free without a card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account and get 1,000 screenshots each month with no card.

FAQ

Is an MCP router part of the MCP specification?

No. MCP defines peer roles and message behavior. Combining several clients behind one server is an architecture built with the SDK APIs.

Should every router forward resources and prompts?

No. Expose only the primitives your host needs, and document how authorization and naming work for each primitive.

Can I use SSE for a new router?

Use Streamable HTTP for new deployed systems. Keep SSE only when a backend still requires compatibility.

How do I avoid exposing every downstream tool to the model?

Filter the discovered catalog by backend, tool name, risk classification, caller identity, or configuration before publishing it.

What should happen when a backend goes down?

Keep a clearly marked last known-good catalog or remove only that backend’s tools, return explicit errors for calls, and expose health information for operators.

Which Python version should I use?

Use Python 3.10 or newer for the current MCP Python SDK v2 line, and pin the SDK major version used by your application.