ScreenshotNeo

BlogAI agents

How to Connect a Local Ollama LLM to an MCP Server

Connect a local Ollama model to MCP tools with a host bridge, runnable Node.js code, transports, troubleshooting, and production guidance.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: Ollama does not plug directly into an MCP server. Ollama provides a chat API that accepts tool definitions and returns tool calls. An MCP-capable host or client must discover tools from the MCP server, translate their schemas into Ollama’s tool format, execute requested calls through MCP, and send the results back to Ollama.

The bridge can run entirely on your machine. Ollama handles local inference; the MCP client handles protocol discovery and tool execution. A tool can still access the network or local files, so inspect each MCP server before treating the complete workflow as private.

How the connection works

  1. Start Ollama and load a model whose current documentation lists tool-calling support.
  2. Connect an MCP client to a server over stdio or Streamable HTTP.
  3. Call MCP listTools and convert each tool’s name, description, and JSON input schema to Ollama’s tools array.
  4. Send the user’s conversation and those tools to Ollama’s chat endpoint.
  5. If Ollama returns a tool call, validate its name and arguments, invoke the matching MCP tool, and append the result as a tool message.
  6. Send the updated conversation to Ollama again. Stop when the model returns normal text or your iteration and time limits are reached.

This orchestration follows the documented capabilities of the Ollama and MCP APIs; Ollama’s model endpoint is not itself an MCP client. Ollama documents tool definitions and tool_calls in its tool support documentation, while the MCP TypeScript SDK documents client transports and tool operations.

Prerequisites

  • Ollama installed and running locally.
  • A pulled model tag that supports tool calling. Support changes by model and tag, so check the current model documentation.
  • An MCP server and its required command, arguments, environment variables, or HTTP endpoint.
  • Node.js 18 or newer for the example below.

Ollama’s official streaming article says that a context window of 32k or higher may improve tool calling anecdotally, while longer contexts use more memory. Treat that as guidance to measure, not a universal minimum.

Complete Node.js bridge

The following host uses the MCP TypeScript SDK with a locally spawned server over stdio and Ollama’s HTTP API. Replace the command and model with values for your server and model. Consult the ScreenshotNeo documentation for the same MCP concepts when using ScreenshotNeo’s MCP server.

mkdir ollama-mcp-bridge
cd ollama-mcp-bridge
npm init -y
npm install @modelcontextprotocol/sdk

Create bridge.mjs:

import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';

const OLLAMA_URL = process.env.OLLAMA_URL ?? 'http://127.0.0.1:11434';
const MODEL = process.env.OLLAMA_MODEL ?? 'llama3.1';

// Change these to the MCP server's documented startup command and arguments.
const transport = new StdioClientTransport({
  command: process.env.MCP_COMMAND ?? 'node',
  args: (process.env.MCP_ARGS ?? './server.mjs').split(' ').filter(Boolean),
  env: { ...process.env }
});

const mcp = new Client({ name: 'ollama-mcp-bridge', version: '1.0.0' }, { capabilities: {} });
await mcp.connect(transport);

const discovered = await mcp.listTools();
const tools = discovered.tools.map((tool) => ({
  type: 'function',
  function: {
    name: tool.name,
    description: tool.description ?? '',
    parameters: tool.inputSchema ?? { type: 'object', properties: {} }
  }
}));

const messages = [
  { role: 'user', content: process.argv.slice(2).join(' ') || 'Describe the available tools.' }
];

for (let turn = 0; turn < 8; turn += 1) {
  const response = await fetch(`${OLLAMA_URL}/api/chat`, {
    method: 'POST',
    headers: { 'content-type': 'application/json' },
    body: JSON.stringify({ model: MODEL, messages, tools, stream: false })
  });

  if (!response.ok) throw new Error(`Ollama HTTP ${response.status}: ${await response.text()}`);
  const data = await response.json();
  const assistant = data.message;
  messages.push(assistant);

  const calls = assistant.tool_calls ?? [];
  if (calls.length === 0) {
    console.log(assistant.content ?? '');
    break;
  }

  for (const call of calls) {
    const name = call.function?.name;
    const args = call.function?.arguments ?? {};
    const known = discovered.tools.find((tool) => tool.name === name);
    if (!known) throw new Error(`Ollama requested an undiscovered tool: ${name}`);

    let result;
    try {
      result = await mcp.callTool({ name, arguments: args });
    } catch (error) {
      result = { isError: true, content: [{ type: 'text', text: String(error) }] };
    }

    messages.push({
      role: 'tool',
      content: JSON.stringify(result)
    });
  }
}

await mcp.close();

Run it with a local server command:

OLLAMA_MODEL=your-tool-capable-model \
MCP_COMMAND=node \
MCP_ARGS="./server.mjs" \
node bridge.mjs "Use the read-only tool to summarize the current data"

The SDK API can change. Pin a compatible SDK version in your application, review its current connection example, and adjust imports or result formatting when upgrading.

Using Streamable HTTP instead of stdio

Use stdio when the host starts a local child process. Use Streamable HTTP when the MCP server exposes a reachable endpoint. The MCP transport documentation also specifies protocol-version handling for subsequent HTTP requests. Configure authentication according to the server’s documentation.

import { Client } from '@modelcontextprotocol/sdk/client/index.js';
import { StreamableHTTPClientTransport } from '@modelcontextprotocol/sdk/client/streamableHttp.js';

const mcp = new Client({ name: 'ollama-http-bridge', version: '1.0.0' }, { capabilities: {} });
const transport = new StreamableHTTPClientTransport(
  new URL(process.env.MCP_URL),
  { requestInit: { headers: { Authorization: `Bearer ${process.env.MCP_TOKEN}` } } }
);
await mcp.connect(transport);
const { tools } = await mcp.listTools();
// Translate tools, call Ollama, then use mcp.callTool({ name, arguments }).

Do not make an old SSE recipe your default. The TypeScript client documentation identifies Streamable HTTP for HTTP servers and describes SSE as a fallback for servers that only support SSE.

Ollama API examples

A direct chat request proves that the model can receive tools. It does not perform MCP discovery or execution by itself.

curl http://127.0.0.1:11434/api/chat \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "your-tool-capable-model",
    "stream": false,
    "messages": [{"role":"user","content":"What can you do?"}],
    "tools": [{
      "type":"function",
      "function": {
        "name":"get_weather",
        "description":"Get weather for a city",
        "parameters": {
          "type":"object",
          "properties":{"city":{"type":"string"}},
          "required":["city"]
        }
      }
    }]
  }'
import requests

payload = {
    "model": "your-tool-capable-model",
    "stream": False,
    "messages": [{"role": "user", "content": "What can you do?"}],
    "tools": [{"type": "function", "function": {
        "name": "get_weather",
        "description": "Get weather for a city",
        "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}
    }}]
}
r = requests.post("http://127.0.0.1:11434/api/chat", json=payload, timeout=120)
r.raise_for_status()
print(r.json())
const response = await fetch('http://127.0.0.1:11434/api/chat', {
  method: 'POST',
  headers: {'content-type': 'application/json'},
  body: JSON.stringify({
    model: 'your-tool-capable-model',
    stream: false,
    messages: [{role: 'user', content: 'What can you do?'}],
    tools: [{type: 'function', function: {
      name: 'get_weather', description: 'Get weather for a city',
      parameters: {type: 'object', properties: {city: {type: 'string'}}, required: ['city']}
    }}]
  })
});
console.log(await response.json());

Schema translation and result handling

  • Preserve the MCP tool name exactly, or maintain a deterministic name map.
  • Copy descriptions that help the model choose safely.
  • Pass the MCP input schema as Ollama function parameters, preserving types, required fields, enums, and nested objects.
  • Validate model-generated arguments before dispatch. Never assume a model’s JSON is safe or complete.
  • Return structured MCP results as text or JSON that the model can understand, including an explicit error when isError is true.
  • Allow multiple tool calls only when your host can execute them safely; otherwise process one call per turn.
  • Set maximum turns, request timeouts, and cancellation behavior to prevent loops.

Security and privacy checklist

  • Allow-list tools instead of exposing every discovered capability.
  • Require confirmation for writes, deletes, payments, messages, or shell commands.
  • Validate URLs, paths, headers, and identifiers at the host boundary.
  • Keep MCP secrets in environment variables or a secret manager.
  • Log tool name, duration, and success state without logging tokens or sensitive arguments.
  • Review whether the server sends data to external services or reads local files.

Troubleshooting

Symptom Likely cause Fix
No tool call appears The model does not support tools, tools were omitted, or only assistant text was inspected. Check the exact model tag, include the tools field, use stream:false while debugging, and inspect message.tool_calls.
Unknown tool name The model returned a name that was not discovered or a translation changed it. Keep a name map, reject unknown names, and send the discovered schemas again.
Invalid arguments The schema was incomplete or the model produced the wrong type. Preserve the MCP JSON schema and validate arguments before callTool.
stdio server will not start Wrong command, arguments, working directory, permissions, or environment. Run the command manually, use absolute paths, print stderr, and compare the host configuration with the server’s instructions.
HTTP MCP connection fails Wrong endpoint, authentication, transport, or protocol-version handling. Verify the Streamable HTTP URL and headers and use a current compatible SDK.
Responses degrade with long conversations Longer context consumes more memory and may exceed the model’s context. Trim old messages, summarize tool output, and measure a suitable context size. Ollama’s 32k+ suggestion is anecdotal.
Tool loop never ends No maximum iteration count or the model keeps requesting another call. Set a turn limit, detect repeated calls, and return a clear stop message.

Performance, reliability, and cost

Local inference avoids a hosted model request, but latency still depends on model size, available memory, context length, tool execution, and network calls made by the MCP server. Measure the complete loop rather than Ollama generation alone.

  • Discover tools once per session and cache schemas until the server changes.
  • Keep tool descriptions and returned payloads concise.
  • Use bounded timeouts for both Ollama and MCP calls.
  • Retry only idempotent operations, with backoff.
  • Persist conversation state if a process may restart, but do not persist secrets unnecessarily.
  • Use a read-only tool for the first integration test.

Ollama itself is local software. MCP tools may incur their own service charges or use external APIs. There is no universal hardware or latency requirement for every model; choose and measure against your machine and model tag.

Or skip the browser setup

If your MCP agent needs website screenshots, ScreenshotNeo provides an MCP server with take_screenshot, get_page_info, and capture_pdf. You can connect that server through the same host bridge, or call its API directly:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing state. Its MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free without a card; paid plans start at $5 for 3,000. See the API and MCP documentation, then sign up free.

FAQ

Can Ollama use MCP tools?

Yes, through an MCP-capable host or bridge that translates discovered MCP tools into Ollama tool definitions and dispatches calls back to MCP.

Does Ollama’s API connect to an MCP URL directly?

No. The host application must create the MCP client connection and perform discovery and execution.

Which transport should I choose?

Choose stdio for a locally spawned server and Streamable HTTP for a server endpoint. Follow the specific server’s authentication and version requirements.

Can every Ollama model call tools?

No. Check the current documentation for the exact model tag and verify its returned message structure.

Is the whole workflow private?

Inference can stay on your local Ollama runtime, but an MCP server may access files or external services. Review its implementation and permissions.

Primary sources