How to Build Infrastructure for AI Agents
A practical architecture for production AI agents: models, tools, memory, runtimes, security, observability, evaluation, and cost controls.
Direct answer: Build an AI agent as a layered software system, not as a prompt wrapped around a model API. Start with one agent and a deterministic workflow. Add model access, authorized tools, retrieval, external state, a controlled runtime, identity and policy, observability, evaluation, and cost limits. Introduce multiple agents only when specialization, parallel work, or separate security domains justify the coordination overhead.
1. The infrastructure an AI agent needs
A production stack normally includes these layers. AWS separates model access, tools, knowledge bases, memory, and orchestration into distinct services, while Google describes frontend, framework, tools, memory, patterns, runtime, models, and model runtime as core components.
| Layer | What it provides | Key decisions |
|---|---|---|
| User/application | UI or API, authentication, streaming responses, sessions | Internal demo or external product; synchronous or streaming |
| Agent logic | Instructions, planning, routing, handoffs, deterministic steps | Testability, framework control, failure handling |
| Model access | Foundation-model APIs, routing, guardrails, quotas, cost allocation | Quality, latency, price, residency, fallback |
| Tools and protocols | Functions, APIs, MCP servers, code execution, agent-to-agent calls | Authorization, validation, timeout, retry, blast radius |
| Knowledge and memory | Approved-data retrieval, session state, durable memory | Freshness, access control, durability, recall quality |
| Runtime | Managed host, container platform, or Kubernetes | Language, portability, isolation, scaling, customization |
| Operations | Logs, traces, evaluations, alerts, release controls | Debuggability, regression detection, auditability, cost visibility |
| Governance and security | Identity, least privilege, policy, approvals, data boundaries | Risk tier, compliance, accountability |
See AWS’s enterprise agentic AI architecture and Google’s component selection guide for the source models.
2. Write the agent charter before writing code
Create a short, versioned charter. Microsoft describes it as the authoritative reference for what the system accomplishes and must avoid.
- Purpose and supported users
- Allowed actions and prohibited actions
- Data the agent may read, write, retain, or transmit
- Tools it may call and the conditions for each call
- Human escalation points and approval requirements
- Success criteria, latency target, and maximum cost per task
- Failure behavior: stop, retry, ask a question, or escalate
Turn the charter into policy tests. A prompt is not an authorization boundary.
3. Start with a deterministic single-agent workflow
Google calls a single-agent system an effective starting point. Use explicit steps when accountability and debugging matter:
- Validate and classify the request.
- Retrieve only the data needed for this task.
- Ask the model for a structured plan or decision.
- Authorize each proposed tool call.
- Validate arguments against a schema.
- Execute with deadlines and idempotency keys.
- Validate the result and record an audit event.
- Return a response or escalate to a human.
Use parallel branches only after you can explain coordination, partial failure, retries, and cancellation. Multi-agent systems add evaluation, security, and operational overhead.
Minimal Python orchestration skeleton
from dataclasses import dataclass
from typing import Any
@dataclass
class ToolResult:
ok: bool
value: Any = None
error: str | None = None
def run_agent(request, model, retriever, authorize, tools, audit):
audit("request.received", {"request_id": request["id"]})
context = retriever.search(request["text"], user=request["user"])
plan = model.plan({"request": request["text"], "context": context})
for call in plan.get("tool_calls", []):
if not authorize(request["user"], call["name"], call["arguments"]):
audit("tool.denied", {"tool": call["name"]})
return {"status": "needs_approval"}
result: ToolResult = tools[call["name"]](**call["arguments"])
audit("tool.completed", {"tool": call["name"], "ok": result.ok})
if not result.ok:
return {"status": "failed", "error": result.error}
answer = model.respond({"request": request["text"], "plan": plan})
audit("request.completed", {"request_id": request["id"]})
return {"status": "ok", "answer": answer}
Keep orchestration state explicit so a run can be resumed, inspected, or replayed.
4. Model access and routing
Put model calls behind one internal interface. Enforce allowed models, data handling rules, token limits, rate limits, retries, fallbacks, and per-tenant cost attribution there. Record model name, request ID, input and output token counts, latency, safety decisions, and errors.
Choose models by task rather than by brand: a smaller model may classify or extract fields, while a stronger model handles ambiguous planning. Define a fallback policy that preserves the output schema and does not silently reduce safety controls.
5. Tools are security boundaries
Every API, MCP server, database, SaaS connector, or code executor is a privileged capability. AWS identifies inbound and outbound authentication and authorization as separate concerns.
- Use a workload identity for the agent and task-scoped credentials for tools.
- Store secrets in a secret manager; never place them in prompts, logs, or source control.
- Allow-list tool names and endpoints.
- Validate arguments with strict schemas, ranges, ownership checks, and tenant filters.
- Set timeouts, retry budgets, payload limits, and circuit breakers.
- Make writes idempotent and require approval for irreversible or high-impact actions.
- Validate tool results before passing them back to the model.
- Log who authorized a call, which policy matched, and what happened.
For browser automation, isolate sessions and treat downloaded files, web content, and instructions returned by a site as untrusted input.
6. Knowledge retrieval and memory
Separate knowledge from memory. Retrieval answers questions from approved, indexed sources. Memory stores information about a session or user that the agent is permitted to retain.
Retrieval checklist
- Ingest documents with source, owner, tenant, classification, and version metadata.
- Filter by authorization before similarity search results reach the model.
- Return citations or source IDs with every retrieved passage.
- Re-index when source content changes and remove deleted material.
- Measure recall and answer faithfulness with a fixed evaluation set.
Short-term and long-term memory
Keep the current turn and workflow state in session memory. Store durable preferences or facts only with an explicit retention policy, deletion path, and access check. Google recommends external persistent storage for production applications because in-memory state disappears when instances terminate.
CREATE TABLE agent_runs (
run_id TEXT PRIMARY KEY,
tenant_id TEXT NOT NULL,
status TEXT NOT NULL,
state JSONB NOT NULL,
created_at TIMESTAMP NOT NULL,
updated_at TIMESTAMP NOT NULL
);
CREATE INDEX agent_runs_tenant_status ON agent_runs (tenant_id, status);
7. Pick a runtime
| Runtime | Use it when | Trade-off |
|---|---|---|
| Managed Agent Runtime | You want built-in lifecycle, scaling, memory, identity, and observability | Opinionated environment and less low-level control |
| Cloud Run | You want containerized stateless services, custom tools, and scale-to-zero | Attach external stores for state; manage more components yourself |
| GKE | You need Kubernetes-level control, complex topology, or existing GKE operations | Higher platform and operations overhead |
| Bedrock AgentCore | AWS-native runtime, MCP gateway, memory, identity, observability, evaluations, and Cedar policy fit your stack | More dependence on AWS services and controls |
Google documents the managed runtime, Cloud Run, and GKE patterns in its architecture guide. Select based on control versus speed, portability, isolation, data residency, latency, and operating cost.
8. Deployment pattern
- Package the agent and tools as an immutable container.
- Inject configuration and secrets at runtime.
- Use a queue for long-running or retryable jobs; return a run ID to the caller.
- Persist checkpoints after retrieval, planning, and each tool call.
- Use separate identities and data stores for development, staging, and production.
- Release prompts, tool schemas, policies, and code together with a version ID.
- Roll out gradually and keep a rollback version.
Keep synchronous requests short. Streaming improves perceived latency, while asynchronous jobs protect the API from model or tool calls that exceed request timeouts.
9. Observability and evaluation
Normal CPU and memory metrics are not enough. AWS recommends tracing model calls, tool selections and results, failures, policy events, quality evaluations, latency, and cost.
| Signal | Record |
|---|---|
| Trace | Run, model call, retrieval, policy decision, tool call, retry, handoff |
| Quality | Task success, schema validity, groundedness, escalation rate, human rating |
| Reliability | Timeouts, tool error rate, queue age, incomplete runs, retry count |
| Cost | Tokens, model price, tool charges, storage, egress, cost per successful task |
| Safety | Denied actions, prompt-injection detections, sensitive-data events, approvals |
Build a regression set before changing prompts or models. Replay representative requests, compare structured outputs and policy events, and require human review for high-impact domains.
10. Performance, reliability, and cost controls
- Reduce latency: stream responses, cache immutable retrieval results, parallelize independent reads, and avoid unnecessary model turns.
- Bound work: cap steps, tool calls, tokens, wall-clock time, retries, and recursion depth.
- Design for partial failure: use deadlines, exponential backoff with jitter, circuit breakers, and resumable checkpoints.
- Control spend: route simple tasks to cheaper models, enforce per-user and per-run budgets, cache safely, and stop runs that exceed their plan.
- Protect availability: queue bursts, apply backpressure, and keep a degraded response path when retrieval or a nonessential tool is unavailable.
- Preserve correctness: prefer deterministic workflow steps for critical business logic and require confirmation before irreversible writes.
Agent requests can trigger several inference calls, tool invocations, memory lookups, and inter-agent communications; each adds latency, cost, and failure surface. See the AWS Agentic AI Lens.
11. Browser screenshots as an agent tool
If an agent needs visual evidence from a web page, define a narrow screenshot tool with an allow-listed URL policy, timeout, maximum image size, and retention rule. Capture only the page or element needed for the task, and treat page content as untrusted.
DIY browser tool considerations
- Launch an isolated browser context per tenant or job.
- Set a fixed viewport, device scale, locale, timezone, and user agent.
- Wait for a selector or network idle instead of sleeping indefinitely.
- Block unnecessary ads, trackers, and resource types to reduce cost and latency.
- Hide sensitive selectors and avoid storing screenshots longer than required.
- Classify bot checks, blank pages, timeouts, and failed loads as explicit outcomes.
12. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Use the API directly from an agent tool:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the complete parameter list. It supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF paper size and page ranges, HTML/CSS to image, custom CSS and JavaScript, clicks, selector waits, delays, network idle, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Existing screenshot API parameter names also work, which simplifies migration.
ScreenshotNeo has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
13. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| Agent repeats a tool call | No idempotency key or unclear result state | Persist call status, use an idempotency key, and cap retries |
| Stale or unauthorized answers | Retrieval runs before access filtering or indexing is old | Filter by tenant and role before search; re-index changed sources |
| State disappears after deployment | State stored only in process memory | Persist checkpoints and memory in an external durable store |
| Requests time out | Too many sequential model or tool calls | Set deadlines, parallelize independent work, queue long jobs, and reduce context |
| Unexpected spend | Unbounded loops, retries, or token budgets | Set per-run budgets, maximum steps, model routing, and cost alerts |
| Unsafe action executes | Tool authorization delegated to the model | Enforce policy in code, validate arguments, and require human approval |
| Screenshot contains a popup | Consent or widget cleanup was disabled or unsupported | Enable cleanup, add a hide selector, or use custom CSS |
| Screenshot is blank or blocked | Bot check, failed load, or page timeout | Inspect verdict headers, adjust waits or headers, and treat the result as a non-success outcome |
14. When to introduce multiple agents
Split into agents only when a boundary provides measurable value: distinct expertise, parallel work with independent failure handling, separate credentials or data domains, or an independently deployable team. Define the coordinator’s contract, message schema, timeout, retry policy, and authority for every handoff. Keep a single-agent fallback for routine requests.
15. Production readiness checklist
- Agent charter is versioned and reviewed.
- Critical logic is deterministic and replayable.
- Every tool has an owner, schema, authorization policy, timeout, and audit event.
- Secrets, identity, tenant isolation, and data retention are enforced outside the prompt.
- Session state and checkpoints survive instance termination.
- Model, tool, retrieval, safety, latency, and cost traces are searchable.
- Regression evaluations run before releases.
- Budgets, rate limits, backpressure, and escalation paths are configured.
- Runbooks cover model outage, tool outage, stale data, unsafe output, and rollback.
FAQ
Should I use an agent framework?
Use one when it reduces orchestration code without hiding authorization, state, tracing, or failure behavior. A small explicit workflow is often easier to operate first.
Is a vector database required?
No. Use retrieval only when the agent needs changing or private source material. A relational store, search service, or document index can be sufficient.
Should every response be streamed?
Stream interactive text when it improves the user experience. Keep side effects and final status behind explicit, auditable workflow steps.
How do I choose between managed and self-managed hosting?
Choose managed hosting for faster delivery and built-in operations; choose containers or Kubernetes when portability, isolation, topology, or low-level control outweigh maintenance cost.
What is the first production metric?
Track successful task completion with latency and cost per successful task. Add safety, tool, retrieval, and quality dimensions immediately after.
The research framing in Infrastructure for AI Agents describes infrastructure as the systems and protocols that attribute actions, shape interactions, and detect or remedy harmful behavior. Those responsibilities should be visible in your architecture from the first release.


