ScreenshotNeo

BlogAI agents

AI Gateway: Definition and How It Works

An AI gateway gives applications one controlled interface for many model providers, handling routing, security, policy, observability, and cost.

By the ScreenshotNeo team1 October 20267 min read

An AI gateway is a software intermediary between an application, agent, or MCP client and one or more AI model providers. It presents a stable interface to callers while handling provider-specific authentication, request translation, routing, policy enforcement, retries, observability, and usage accounting.

The gateway becomes the controlled boundary between your software and providers such as OpenAI, Anthropic, Azure OpenAI, Amazon Bedrock, Google Gemini, or self-hosted models. Clients send one normalized request; the gateway chooses an upstream target, applies credentials and policies, translates the protocol, forwards the call, and returns a normalized response.

What an AI gateway does

Capability What it solves
Provider abstraction Applications use one contract instead of embedding every provider’s API format and authentication method.
Routing and failover Requests can be sent by model, priority, latency, cost, usage, or capacity, with retries and circuit breaking where configured.
Credential management Provider keys, cloud signatures, managed identities, or service accounts stay at the gateway boundary.
Governance Authentication, authorization, rate limits, ACLs, data policies, and guardrails apply consistently across teams.
Observability Request counts, latency, errors, tokens, and estimated cost can be collected in one place.
Protocol coverage Gateways may proxy chat, embeddings, image, audio, video, realtime, MCP, and agent-to-agent traffic.

How an AI gateway works

  1. Client sends a request. An application, orchestration service, MCP client, or agent calls the gateway endpoint.
  2. The gateway selects a target. A model name or policy maps the request to one or more configured providers.
  3. Credentials are attached. The gateway adds the correct API key, cloud signature, managed identity, or service account.
  4. Formats are translated. A normalized request is converted to the provider’s native protocol; the response is converted back.
  5. Policies run. Authentication, authorization, rate controls, content rules, and data-governance checks are evaluated.
  6. The call is forwarded and observed. The gateway records configured telemetry and may retry or fail over.
  7. A normalized response returns. The caller receives a consistent shape even when the provider changes.

Kong describes this mediation explicitly: at request time, the AI model mediates traffic between clients and upstream AI provider APIs. The same pattern applies to managed gateways, self-hosted proxies, and hybrid deployments.

Reference architecture

A production design commonly separates a control plane from one or more data planes.

  • Control plane: stores model targets, credentials, policies, routing rules, certificates, and configuration distribution.
  • Data plane: receives live requests, enforces the distributed configuration, calls providers, and emits telemetry.
  • Provider layer: contains commercial APIs, cloud model services, or self-hosted inference endpoints.
  • Observability layer: collects latency, status, token, cost, and optional payload data.

In a hybrid topology, the control plane can be managed while data-plane nodes run in your network. This keeps live traffic closer to private systems, but you still own capacity planning, upgrades, secrets handling, high availability, and monitoring for those nodes.

AI gateway versus API gateway

Area Traditional API gateway AI gateway
Primary traffic REST, GraphQL, and service APIs LLM, embedding, multimodal, realtime, MCP, and agent traffic
Routing key Host, path, method, or service Model, provider, latency, token budget, cost, quota, or semantic policy
Usage data Requests, bytes, status, latency Those metrics plus tokens, model, estimated cost, and provider usage
Translation HTTP transformations and authentication Provider-specific message, tool, streaming, and response formats
Policy concerns Identity, access, quotas, and network security Those controls plus prompts, outputs, sensitive data, model allow-lists, and retention

An AI gateway is often an API gateway with model-aware routing, translation, and accounting. Existing API-gateway products can provide AI features; dedicated AI gateways generally expose deeper provider and token controls.

When you need one

  • Several applications use more than one model provider.
  • Provider credentials must be removed from application code.
  • You need centralized quotas, tenant isolation, or spend attribution.
  • Provider outages require automatic fallback.
  • Teams need one audit and observability surface.
  • You are standardizing MCP or agent-to-agent traffic.

A gateway may be unnecessary for a small application that calls one provider directly. It adds another network hop, configuration surface, dependency, and possible privacy boundary. Adopt it when those costs are smaller than the control and portability it provides.

Routing, retries, and failover

Common routing strategies include round robin, consistent hashing, least connections, lowest latency, lowest usage, semantic routing, and explicit priority. A robust policy defines:

  • Which models and providers are allowed for each workload.
  • What errors are retryable (usually timeouts, connection failures, and selected 5xx responses).
  • How many attempts are permitted and the maximum retry delay.
  • Whether a fallback may change model quality, region, or data residency.
  • When a circuit opens after repeated failures and how it closes.

Use bounded exponential backoff with jitter. Never retry non-idempotent tool calls blindly; duplicate side effects can be more damaging than a failed request.

Security and governance checklist

  • Keep provider secrets in a secret manager or gateway credential store.
  • Authenticate callers with short-lived tokens or workload identity.
  • Authorize by application, tenant, model, tool, and environment.
  • Apply per-user and global rate limits before forwarding traffic.
  • Allow-list providers, regions, models, and tool servers.
  • Define prompt, output, and sensitive-data logging rules before enabling payload logs.
  • Encrypt traffic in transit and protect gateway-to-provider certificates.
  • Record policy decisions and configuration changes for audit.

Streaming and protocol support

Verify support for server-sent events, HTTP/2, WebSockets, tool calls, and provider-specific streaming semantics. A gateway that only handles buffered JSON may break token streaming or realtime audio. Test cancellation, client disconnects, partial responses, and upstream timeouts under load.

Deployment choices

Model Advantages Trade-offs
Managed Fast setup, vendor-operated scaling and upgrades Less network control and an additional vendor for configuration or telemetry
Self-hosted Maximum placement, customization, and data-path control You operate capacity, upgrades, secrets, HA, and incident response
Hybrid Managed configuration with data planes in your network More moving parts and certificate/configuration distribution to manage

Minimal normalized request example

The exact endpoint and fields depend on the gateway. This generic shape shows the contract your application should own rather than provider-specific details:

curl https://gateway.example.com/v1/chat/completions \
  -H 'Authorization: Bearer GATEWAY_TOKEN' \
  -H 'Content-Type: application/json' \
  -d '{"model":"production-chat","messages":[{"role":"user","content":"Summarize this incident."}],"stream":false}'

Python

import requests

response = requests.post(
    "https://gateway.example.com/v1/chat/completions",
    headers={"Authorization": "Bearer GATEWAY_TOKEN"},
    json={
        "model": "production-chat",
        "messages": [{"role": "user", "content": "Summarize this incident."}],
    },
    timeout=60,
)
response.raise_for_status()
print(response.json())

Node.js

const response = await fetch('https://gateway.example.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    authorization: 'Bearer GATEWAY_TOKEN',
    'content-type': 'application/json'
  },
  body: JSON.stringify({
    model: 'production-chat',
    messages: [{ role: 'user', content: 'Summarize this incident.' }]
  })
});
if (!response.ok) throw new Error(`${response.status} ${await response.text()}`);
console.log(await response.json());

Performance, reliability, and cost

  • Latency: measure gateway overhead separately from provider time. Connection reuse, regional placement, and streaming reduce perceived delay.
  • Capacity: scale data-plane workers for concurrent streams, not only request rate. Protect providers with queues and admission limits.
  • Reliability: monitor gateway health, provider health, retry volume, fallback rate, timeout rate, and circuit state.
  • Cost: account for gateway hosting or subscription fees, provider tokens, retries, logging storage, egress, and duplicated traffic during failover.
  • Privacy: determine whether prompts, outputs, metadata, or telemetry leave your network and how long they are retained.

No cross-vendor performance or price benchmark is implied here. Provider coverage and gateway features change, so verify current documentation before choosing an implementation.

Troubleshooting

Symptom Likely cause Fix
401 or 403 Invalid gateway token, missing provider credential, or policy denial Check caller identity, credential mapping, model allow-lists, and gateway audit logs.
404 model Model alias is not mapped to an upstream target Register the alias and confirm the target is enabled in the deployed configuration.
429 responses Gateway or provider quota exceeded Inspect which limit fired, add bounded backoff, and provision quota before raising concurrency.
Timeouts Provider latency, queue saturation, or an overly short gateway timeout Measure each hop, reuse connections, tune deadlines, and cap retries.
Streaming stalls Buffering proxy, unsupported protocol, or idle timeout Enable SSE/WebSockets passthrough, disable buffering, and test heartbeat and cancellation behavior.
Unexpected provider bill Retries, fallback, or payload logging increased usage Track attempts and tokens per request, then set budgets and retry caps.
Data appears in logs Payload logging enabled by default Disable body logging or redact sensitive fields and set a short retention period.

How to evaluate an AI gateway

  1. List providers, models, protocols, regions, and compliance requirements.
  2. Define a normalized request and response contract owned by your application.
  3. Test routing, retries, fallback, streaming, cancellation, and malformed responses.
  4. Measure added latency and cost at realistic concurrency.
  5. Review secret storage, tenant isolation, logs, retention, and operator access.
  6. Run an outage exercise with one provider unavailable.
  7. Choose managed, hybrid, or self-hosted deployment based on operational capacity.

Or skip the browser setup

When an AI agent needs screenshots of documentation, dashboards, or test pages, ScreenshotNeo provides a website screenshot API and MCP server. A single request returns PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, cookie banners, newsletter popups, and chat widgets are removed. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. See the ScreenshotNeo API documentation, then sign up free.

FAQ

Does an AI gateway replace a model provider?

No. It governs and translates traffic to providers; it does not supply model inference by itself.

Can one gateway expose non-LLM tools?

Many current gateways also proxy MCP tool servers and agent-to-agent traffic, subject to product support.

Should every request be retried?

No. Retry only transient failures within a bounded budget, and avoid blind retries for side-effecting tool calls.

Is a gateway always cheaper?

No. It can reduce waste through routing and accounting, but adds infrastructure or subscription cost and may increase latency.

How often should provider mappings change?

Review them whenever providers change models, pricing, limits, regions, or protocol behavior, and validate mappings in a staging environment first.