Cloudflare Workers MCP Server
Learn the three meanings of Cloudflare Workers MCP Server, build a remote Streamable HTTP server, test it, secure it, and deploy it.
Cloudflare Workers MCP Server can mean three different things: the older workers-mcp package that bridges a local MCP client to a Worker, a custom remote MCP server that you deploy on Workers, or Cloudflare-hosted MCP servers that let agents operate Cloudflare products. Choose the option that matches your goal before writing code.
For a new integration, the most direct path is usually a custom remote server using Streamable HTTP. Cloudflare’s current guide covers local testing with Wrangler and the MCP Inspector, then deployment with Wrangler to a /mcp endpoint. Commands and package APIs can change, so check the current Cloudflare remote MCP guide before publishing a production deployment.
Which Cloudflare Workers MCP server do you mean?
| Option | Use it when | Where it runs | Tool scope |
|---|---|---|---|
workers-mcp package |
You already have TypeScript methods in a Worker and want build tooling plus a local bridge. | Worker in Cloudflare; local Node.js process proxies MCP stdio. | Your Worker methods translated into MCP tools. |
| Custom remote MCP server | You are exposing your own service to remote AI clients. | A Worker reachable over Streamable HTTP, commonly at /mcp. |
Tools you define. |
| Cloudflare-hosted MCP servers | You want an agent to work with Cloudflare APIs and products. | Cloudflare-operated endpoints. | Code Mode for broad API access or curated product-specific tools. |
The workers-mcp repository README itself points readers toward the remote-server approach for new projects. Cloudflare’s mcp-server-cloudflare repository describes hosted servers, while the cloudflare/mcp repository documents Code Mode and product-specific servers.
Architecture of a remote MCP server on Workers
A remote setup has four pieces:
- An MCP client such as Claude, Cursor, or another MCP-compatible agent.
- A Streamable HTTP connection to your Worker.
- The Worker MCP endpoint, normally mounted at
/mcp. - Your tools, which call APIs, Durable Objects, KV, R2, databases, or other services through bindings.
Unlike a local stdio server, the remote endpoint must handle authentication, concurrent requests, connection lifecycle, and deployment configuration. Decide who may call each tool before making the endpoint public.
Build a custom remote MCP server
1. Create the Worker project
npm create cloudflare@latest my-mcp-server
cd my-mcp-server
npm install
npm install @modelcontextprotocol/sdk
npm install -D wrangler
Select a Worker application when prompted. The exact scaffold questions and package versions are version-sensitive; use the choices in Cloudflare’s current guide.
2. Define a tool and an HTTP entry point
The MCP SDK API changes over time. The following is a compact TypeScript shape for the current SDK family; if an import or transport constructor differs in your installed version, copy the equivalent server and transport setup from Cloudflare’s guide.
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StreamableHTTPServerTransport } from "@modelcontextprotocol/sdk/server/streamableHttp.js";
import { z } from "zod";
const server = new McpServer({
name: "example-workers-mcp",
version: "1.0.0"
});
server.tool(
"get_status",
"Return the status of a named service",
{ service: z.string().min(1) },
async ({ service }) => ({
content: [{ type: "text", text: JSON.stringify({ service, status: "ok" }) }]
})
);
export default {
async fetch(request: Request): Promise<Response> {
const url = new URL(request.url);
if (url.pathname !== "/mcp") {
return new Response("Not found", { status: 404 });
}
// Replace this check with your identity provider or signed-token validation.
const token = request.headers.get("authorization");
if (!token) return new Response("Unauthorized", { status: 401 });
const transport = new StreamableHTTPServerTransport({
sessionIdGenerator: undefined
});
await server.connect(transport);
return transport.handleRequest(request);
}
};
For production, use the transport and session pattern recommended by the SDK version in your lockfile. If a tool needs durable per-session state, place that state in a Durable Object rather than a module global.
3. Configure Wrangler
name = "my-mcp-server"
main = "src/index.ts"
compatibility_date = "2026-09-01"
[observability]
enabled = true
Add bindings only for resources your tools actually use. Keep secrets in Wrangler secrets or your deployment system; do not commit API keys.
4. Run locally
npx wrangler dev
Wrangler starts a local Worker URL. The local runtime uses Miniflare and workerd, the same Worker runtime model used in production. Code execution and bindings are separate: bindings are simulated by default unless you explicitly configure remote resources.
5. Inspect the server with MCP Inspector
npx @modelcontextprotocol/inspector
Point the Inspector at your local /mcp URL, include the authorization header required by your Worker, list tools, and invoke each tool with valid and invalid inputs. Test malformed JSON, missing fields, expired credentials, and upstream API failures.
6. Deploy
npx wrangler@latest deploy
Cloudflare’s guide shows a deployed workers.dev address with an /mcp route. Use your actual hostname in the client configuration; do not assume the example hostname or route from an older guide remains unchanged.
Authentication and authorization
An unauthenticated endpoint can be called by anyone who can reach it. That is appropriate only for deliberately public, read-only tools. For anything that changes data, accesses private information, or incurs cost, authenticate the caller and authorize individual operations.
- Validate the issuer, audience, expiry, and signature of bearer tokens.
- Map identities to an allowlist of tools and resource scopes.
- Reject missing or malformed authorization headers with
401; reject authenticated callers without permission with403. - Apply rate limits and upstream timeouts.
- Log tool name, request ID, principal, duration, and outcome without logging secrets.
Using Cloudflare’s hosted MCP servers
Use Cloudflare’s hosted options when the agent needs Cloudflare API capabilities rather than your own business logic. The repository positions Code Mode for broad access across Cloudflare APIs. Domain-specific servers expose smaller, typed tool sets for particular products, including a Workers Bindings server for storage, AI, and compute primitives.
Choose with these questions:
- Do you need your own functions? Build a custom remote server.
- Do you need broad Cloudflare API coverage? Evaluate Code Mode.
- Do you need curated Workers or product operations? Use the relevant domain-specific server.
- Does the agent need write access? Confirm the server’s authorization model and grant the smallest scope.
Code Mode and token footprint
The Cloudflare mcp repository reports a comparison for 2,594 endpoints/tools: approximately 1,100 tokens for Code Mode, 1,170,523 tokens for native MCP with full schemas, and 244,047 tokens for native MCP with only required-parameter schemas. These are figures published by the repository, not an independent benchmark. They describe that repository’s comparison and should not be generalized to every client or workload.
Local development and binding caveats
Local Worker execution can be faithful while the resources behind bindings are not. Miniflare and workerd run your code locally, but KV, R2, Durable Objects, databases, and service bindings may use simulated resources unless configured otherwise. Remote-resource development can expose real data and introduce latency or access-control concerns.
Cloudflare’s local-development documentation states that Workers AI currently has no local simulation. If a tool calls Workers AI, test against a remote resource in a controlled account and make that dependency explicit in your development checklist.
Testing checklist
- Initialize the MCP session and list tools.
- Invoke every tool with its smallest valid input.
- Verify schema failures are returned clearly.
- Test missing, expired, and insufficient credentials.
- Test upstream 4xx, 5xx, timeout, and rate-limit responses.
- Confirm no secret or personal data appears in logs.
- Run the same cases against local simulated bindings and the deployed Worker where remote resources differ.
- Verify a client reconnects cleanly after a dropped HTTP connection.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
404 on /mcp |
Route is different from the client URL or the Worker was not deployed with the expected entry point. | Check the Worker URL, Wrangler configuration, and route handling. |
401 Unauthorized |
Authorization header is missing, expired, or rejected. | Send the expected bearer token and verify issuer, audience, and clock settings. |
| Inspector cannot connect | Wrong transport, URL, or local port; server may only support stdio. | Use the remote Streamable HTTP endpoint and confirm Wrangler is running. |
| Tool appears but invocation fails | Input does not match the declared schema or an upstream binding failed. | Inspect the tool’s schema, validate inputs, and return a useful error with a request ID. |
| Works locally, fails after deploy | Missing secret, binding, compatibility setting, or remote-resource permission. | Compare deployed bindings and secrets with local configuration and inspect Worker logs. |
| Requests hang | Unbounded upstream call, session handling problem, or a promise that never resolves. | Add explicit timeouts, abort signals, and bounded retries; verify transport lifecycle handling. |
| Unexpected data in local tests | A binding is simulated locally while production uses a remote resource, or vice versa. | Document the binding mode and run an integration test against the intended resource. |
Performance, reliability, and cost considerations
Performance
- Keep tool schemas small and precise so clients spend less context describing unavailable operations.
- Set upstream timeouts and avoid serial calls when independent calls can run concurrently.
- Return concise structured results; let the agent request detail when needed.
- Use caching for safe, repeatable reads and batch API calls where the upstream service supports it.
Reliability
- Make mutating tools idempotent where possible and accept an idempotency key for retries.
- Return stable error categories so clients can distinguish invalid input, authorization, rate limits, and temporary failures.
- Use Durable Objects or another persistent store for state that must survive Worker isolates.
- Monitor latency, error rate, authentication failures, and upstream status separately.
Cost
Workers charges depend on the Cloudflare services and plan used by your deployment. MCP itself does not remove costs from upstream APIs, AI calls, storage, or egress. Measure tool call volume and downstream operations, then apply authentication and rate limits before opening an endpoint to many agents.
Or skip the browser setup
If your MCP tool needs website screenshots, you can call ScreenshotNeo instead of maintaining a browser inside a Worker. Its screenshot API accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing result. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for AI agents.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 screenshots a month free with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Should a new project use workers-mcp?
Use it when you specifically need its build tooling and local stdio bridge. For a remote service that agents call directly, follow the current Streamable HTTP guide.
Can a Workers MCP endpoint be public?
Yes, but public access should be limited to intentionally public operations. Sensitive tools should authenticate and authorize every request.
Does local Wrangler perfectly reproduce production?
Worker code runs in the same runtime model, but bindings may be simulated or remote. Workers AI has no current local simulation, so some production behavior requires controlled integration testing.
Is Code Mode always cheaper?
The Cloudflare repository reports a smaller token footprint for its Code Mode comparison. Token counts are repository-reported figures, not a universal cost or latency guarantee.
Where should I verify commands?
Check Cloudflare’s current remote MCP server guide, the local development documentation, and the relevant GitHub README immediately before deployment.


