Infrastructure for Production AI Agents
A practical guide to runtime, tools, memory, security, observability, deployment, and cost controls for reliable production AI agents.

Production AI agent infrastructure is the set of runtime, orchestration, model and tool connections, knowledge, memory, identity, policy, and operational controls that let an agent do useful work under real workload constraints. A model call alone is not production infrastructure. You need to control what the agent can access, observe what it does, recover when a tool fails, evaluate stochastic behavior, and account for every model and tool action.
AWS’s Well-Architected Agentic AI Lens captures the central question: “can we run agents reliably, securely, and cost-effectively at scale?” The rest of this guide turns that question into an architecture and launch checklist.
What infrastructure does an AI agent need in production?
A useful reference model has six interacting layers:
| Layer | What it does | Production decisions |
|---|---|---|
| Application interface | Receives user, event, or API requests and returns results. | Authentication, rate limits, streaming, request limits, tenancy. |
| Agent runtime and orchestration | Runs the reasoning loop, chooses tools, manages sessions, and coordinates agents. | Isolation, timeouts, retries, concurrency, long-running jobs. |
| Models | Generate plans, tool arguments, summaries, and final responses. | Model routing, safety policy, token budgets, fallback models, cost attribution. |
| Tools and actions | Call APIs, execute code, browse sites, update systems, or send messages. | Discovery, contextual authorization, schemas, idempotency, approval gates. |
| Knowledge and memory | Retrieves enterprise data and preserves useful context across turns. | Access control, retention, deletion, isolation, freshness, retrieval cost. |
| Cross-cutting controls | Identity, policy, observability, evaluation, secrets, and incident response. | Auditable traces, anomaly detection, circuit breakers, dashboards. |
AWS’s enterprise reference architecture describes similar application, agents, and agent-accessed service layers. Its model-access category includes policy, safety, guardrails, and cost tracking; its tools category includes discovery, secure execution, and authorization; and its knowledge category exposes data through controlled retrieval systems. These concerns span every layer rather than belonging to one product.
Why agent infrastructure differs from a normal API
A conventional request often has one predictable path: receive input, call a service, return output. An agent may reason through several model calls, retrieve memory, call multiple tools, ask another agent for help, and retry after an error. Each step can add latency, token cost, and a new failure mode.

- Iterative reasoning: completion time and cost depend on the number of loops, not only input size.
- Autonomy: the system can take actions without a person approving every step.
- Stochastic behavior: the same request can produce different plans or tool choices.
- Distributed coordination: multiple agents and remote tools create partial failures and ordering problems.
- Memory: persistent context introduces privacy, integrity, retention, and storage costs.
Design budgets for loops and actions. Set a maximum number of reasoning steps, tool calls, elapsed time, and tokens per task. A task that exceeds a budget should return a controlled status or request human review instead of running indefinitely.
Reference architecture for a production agent
1. Request and session boundary
Terminate TLS and authenticate the caller before the agent runtime. Attach a tenant, user, request ID, and policy context to the session. Store session state outside the worker when work can outlive one process. Make every downstream call carry the same correlation ID so a trace can connect model calls, retrieval, and tools.
2. Runtime and orchestration
Use a worker model that matches the workload:
- Serverless functions: useful for short, stateless orchestration and lightweight tools. Enforce execution and payload limits.
- Managed agent runtimes: reduce lifecycle work and can provide session, sandbox, and observability integrations. Verify data boundaries and supported tools.
- Containers: fit long-running, stateful, CPU-heavy, or custom-dependency workloads. You own deployment, patching, scaling, and capacity planning.
AWS documents managed runtime, serverless, and container implementation options; Google Cloud documents low-code, managed-code, and custom-code paths. OpenAI’s announcement describes an agent harness that can use a provider-managed sandbox, an organization’s infrastructure, or ecosystem environments, including VPC deployments. These are implementation choices, not a universal ranking.
3. Model gateway
Put model calls behind a gateway or shared client. Enforce allowed models, maximum output tokens, temperature ranges, safety settings, and per-tenant budgets there. Record model name, input and output token counts, latency, retries, and estimated cost. Route simple classification or extraction to a smaller model and reserve larger models for tasks that need them, while measuring task quality with representative evaluations.
4. Tool registry and execution
Give every tool a versioned schema describing arguments, permissions, timeout, side effects, and idempotency behavior. Resolve tools through a registry rather than allowing arbitrary URLs or code. Protocols such as MCP and A2A can support discovery and agent-to-agent communication, but discovery must not imply permission.
Execute tools in an isolated process or sandbox when they handle untrusted input or code. Store credentials in a secrets manager, inject short-lived tokens, and redact secrets from traces. Validate both model-generated arguments and tool responses. Use idempotency keys for payments, ticket creation, provisioning, and other actions that must not be repeated after a retry.
5. Knowledge and memory
Separate durable knowledge from conversational memory. Knowledge retrieval should enforce the caller’s document permissions at query time; filtering after retrieval can leak restricted content. Memory should have an explicit schema, owner, retention period, and deletion path. Protect against prompt-injected instructions stored in documents or prior sessions by treating retrieved text as data and rechecking policy before an action.
Persist only what improves future tasks. Summarize long sessions, cap context size, and expire low-value entries. Track retrieval latency and token expansion because memory can become a recurring cost multiplier.
Security, identity, and human oversight
Define an agent’s scope in policy, independent of its system prompt. Prompts guide behavior; they do not enforce authorization.
- Give each agent a distinct identity. Use workload identity or a service principal that can be audited separately from the end user.
- Apply least privilege. Grant only the tools, records, actions, and environments required for the task.
- Authorize each action in context. Check tenant, user, resource, risk, and current approval state immediately before execution.
- Require approval for high-impact actions. Match approval to reversibility: reading a public page needs none; deleting data, moving money, or sending an external message may require a person.
- Validate inputs and outputs. Constrain URLs, SQL, file paths, recipient addresses, and structured arguments. Filter sensitive output before returning it.
- Add circuit breakers. Stop on repeated failures, unusual tool sequences, policy violations, or budget exhaustion.
Google Cloud documents unique agent identity, centralized registration, managed OAuth connections for delegated access, and a gateway that can inspect tool calls and responses. Those are provider-specific implementations of broader identity and policy requirements. Select equivalents that integrate with your existing identity provider and audit system.
Observability and evaluation
Logging a final answer is insufficient. Capture a trace for the complete workflow:
- Request, tenant, user, and policy version.
- Model calls, prompts or prompt references, token counts, latency, and finish reason.
- Retrieval queries, document IDs, filters, and scores.
- Tool name, validated arguments, response status, duration, retries, and approval decision.
- Budget consumption, fallback path, final outcome, and error classification.
Redact secrets and sensitive content before exporting traces. Keep immutable audit records for high-risk actions, with access restricted to operators who need them.
Use two complementary feedback loops. Traces explain what happened in one run; repeatable evaluations tell you whether a change improves outcomes. Build a dataset of representative tasks, tool failures, adversarial inputs, and permission boundaries. Score task success, factuality, policy compliance, unnecessary actions, latency, and cost. Include tool behavior in the evaluation, because a correct plan can still fail when an API returns an error or an authorization check is missing.
Reliability patterns
Timeouts, retries, and fallbacks
Set separate deadlines for the overall task, each model call, retrieval, and each tool. Retry only transient failures with exponential backoff and jitter. Do not blindly retry non-idempotent actions. When a model or tool is unavailable, return a safe partial result, queue the work for later, or route to a configured fallback. Record the fallback so operators can distinguish degraded operation from success.
Long-running and asynchronous work
For tasks that can run for minutes or hours, create a durable job record and return a job ID. Persist the plan and completed steps, checkpoint after side effects, and resume from the last safe point. Send progress events without exposing internal chain-of-thought. Provide cancellation that propagates to tools and workers.
Isolation and multi-tenancy
Isolate session state, files, caches, and credentials by tenant and agent instance. Never use a shared mutable memory namespace without tenant and user keys. Test concurrent sessions with identical prompts to verify that context cannot cross boundaries.
Performance and cost controls
Measure the whole cognitive pipeline rather than only model latency. Common cost and latency multipliers include repeated planning loops, large retrieved documents, multi-agent handoffs, browser rendering, and tool retries.
| Control | How to apply it |
|---|---|
| Loop budget | Cap reasoning turns and tool calls per task. |
| Token budget | Limit context and output; summarize old messages. |
| Model routing | Use the least expensive model that passes task evaluations. |
| Caching | Cache immutable retrieval and deterministic tool results with a defined TTL. |
| Concurrency | Use queues and per-tenant limits to prevent bursts from exhausting capacity. |
| Cost attribution | Charge model, storage, browser, and tool costs to a task, tenant, and workflow version. |
Optimize only after tracing. A shorter prompt that increases retries can cost more; parallel tool calls can reduce latency while increasing rate-limit pressure. Set alerts for spend, token growth, error rates, and unusual loop depth.
Capturing web evidence as an agent tool
Browser capture is a common production tool for visual regression, research, and evidence collection. A do-it-yourself implementation with Playwright can look like this:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 60000 });
await page.screenshot({ path: 'page.png', fullPage: true });
await browser.close();
In production, wrap this in a narrowly scoped tool. Validate the URL against an allowlist, block private network ranges, limit navigation count and downloaded bytes, set a browser timeout, and run the browser in an isolated worker. Decide whether cookies, authentication headers, JavaScript, third-party requests, and full-page lazy loading are allowed. Store the image with tenant-scoped access and a retention policy. Return structured metadata such as final URL, status, duration, and failure reason to the agent.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server for agents. One GET request returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

See the ScreenshotNeo API documentation for all options. The basic calls are:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For agent infrastructure, useful options include full-page capture with lazy images loaded, CSS-selector element capture, device presets or custom viewports, dark mode, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads and resource types, custom headers and cookies, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs, and a usage API. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
ScreenshotNeo has 1,000 free shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account.
Managed platform or your own infrastructure?
Compare deployment paths against your requirements:
| Question | Managed platform | Customer-controlled runtime |
|---|---|---|
| Control and data boundary | Confirm provider regions, retention, sandbox, and network controls. | You control network, storage, runtime image, and placement. |
| Operations | Less patching and scaling work; follow provider limits. | More control; you own upgrades, capacity, and incidents. |
| Integrations | Fast access to supported models, tools, and identity services. | Integrate any service, protocol, or internal system at engineering cost. |
| Workload fit | Good for standard, bursty workloads that match the service. | Useful for long-running, stateful, specialized CPU/GPU workloads. |
| Cost visibility | Usage pricing is simple to start but requires budget limits. | Infrastructure and engineering costs are explicit but distributed. |
Run a small representative workload through each viable path. Compare trace completeness, authorization integration, recovery behavior, latency, and total cost. The reviewed AWS, Google Cloud, and OpenAI documentation describes capabilities and trade-offs; it does not establish a universally best platform.
Production launch checklist
- Every agent has an identity, owner, purpose, and maximum action scope.
- Tool schemas, permissions, timeouts, retries, and idempotency are documented.
- High-impact actions require approval or an explicit reversible workflow.
- Sessions, memory, files, and caches are tenant-isolated.
- Traces connect requests, model calls, retrieval, tools, approvals, and outcomes.
- Representative evaluations cover success, failure, adversarial input, and policy boundaries.
- Budgets exist for loops, tokens, elapsed time, concurrency, and spend.
- Fallbacks, cancellation, checkpointing, and incident runbooks are exercised.
- Secrets and sensitive trace content are protected and redacted.
- Changes are evaluated against a fixed dataset before rollout.
FAQ
Is an agent runtime the same as an agent platform?
No. A runtime executes orchestration. A platform may also provide identity, tool registration, memory, sandboxes, deployment, observability, and evaluation services.
Should memory be a vector database?
Only when semantic retrieval fits the data. Structured state, relational records, object storage, and short-lived caches may be better for other memory types. Choose by access pattern, authorization, retention, and consistency needs.
How much human oversight is required?
Match approval to impact and reversibility. Low-risk reads can be automatic; irreversible, regulated, financial, or externally visible actions generally need stronger review.
Can deterministic tests validate an agent?
They validate components, but stochastic workflows also need repeatable evaluations across representative tasks, tool failures, policy checks, and outcome quality.
When should an agent become asynchronous?
Use a durable job when work can exceed a request timeout, requires many tools, waits on people, or must resume after a worker failure.


