ScreenshotNeo

BlogAI agents

How to Build Infrastructure for AI Agents

A practical architecture for production AI agents: models, tools, memory, runtimes, security, observability, evaluation, and cost controls.

By the ScreenshotNeo team1 October 202610 min read

Direct answer: Build an AI agent as a layered software system, not as a prompt wrapped around a model API. Start with one agent and a deterministic workflow. Add model access, authorized tools, retrieval, external state, a controlled runtime, identity and policy, observability, evaluation, and cost limits. Introduce multiple agents only when specialization, parallel work, or separate security domains justify the coordination overhead.

1. The infrastructure an AI agent needs

A production stack normally includes these layers. AWS separates model access, tools, knowledge bases, memory, and orchestration into distinct services, while Google describes frontend, framework, tools, memory, patterns, runtime, models, and model runtime as core components.

Layer What it provides Key decisions
User/application UI or API, authentication, streaming responses, sessions Internal demo or external product; synchronous or streaming
Agent logic Instructions, planning, routing, handoffs, deterministic steps Testability, framework control, failure handling
Model access Foundation-model APIs, routing, guardrails, quotas, cost allocation Quality, latency, price, residency, fallback
Tools and protocols Functions, APIs, MCP servers, code execution, agent-to-agent calls Authorization, validation, timeout, retry, blast radius
Knowledge and memory Approved-data retrieval, session state, durable memory Freshness, access control, durability, recall quality
Runtime Managed host, container platform, or Kubernetes Language, portability, isolation, scaling, customization
Operations Logs, traces, evaluations, alerts, release controls Debuggability, regression detection, auditability, cost visibility
Governance and security Identity, least privilege, policy, approvals, data boundaries Risk tier, compliance, accountability

See AWS’s enterprise agentic AI architecture and Google’s component selection guide for the source models.

2. Write the agent charter before writing code

Create a short, versioned charter. Microsoft describes it as the authoritative reference for what the system accomplishes and must avoid.

  • Purpose and supported users
  • Allowed actions and prohibited actions
  • Data the agent may read, write, retain, or transmit
  • Tools it may call and the conditions for each call
  • Human escalation points and approval requirements
  • Success criteria, latency target, and maximum cost per task
  • Failure behavior: stop, retry, ask a question, or escalate

Turn the charter into policy tests. A prompt is not an authorization boundary.

3. Start with a deterministic single-agent workflow

Google calls a single-agent system an effective starting point. Use explicit steps when accountability and debugging matter:

  1. Validate and classify the request.
  2. Retrieve only the data needed for this task.
  3. Ask the model for a structured plan or decision.
  4. Authorize each proposed tool call.
  5. Validate arguments against a schema.
  6. Execute with deadlines and idempotency keys.
  7. Validate the result and record an audit event.
  8. Return a response or escalate to a human.

Use parallel branches only after you can explain coordination, partial failure, retries, and cancellation. Multi-agent systems add evaluation, security, and operational overhead.

Minimal Python orchestration skeleton

from dataclasses import dataclass
from typing import Any

@dataclass
class ToolResult:
    ok: bool
    value: Any = None
    error: str | None = None

def run_agent(request, model, retriever, authorize, tools, audit):
    audit("request.received", {"request_id": request["id"]})
    context = retriever.search(request["text"], user=request["user"])
    plan = model.plan({"request": request["text"], "context": context})

    for call in plan.get("tool_calls", []):
        if not authorize(request["user"], call["name"], call["arguments"]):
            audit("tool.denied", {"tool": call["name"]})
            return {"status": "needs_approval"}
        result: ToolResult = tools[call["name"]](**call["arguments"])
        audit("tool.completed", {"tool": call["name"], "ok": result.ok})
        if not result.ok:
            return {"status": "failed", "error": result.error}

    answer = model.respond({"request": request["text"], "plan": plan})
    audit("request.completed", {"request_id": request["id"]})
    return {"status": "ok", "answer": answer}

Keep orchestration state explicit so a run can be resumed, inspected, or replayed.

4. Model access and routing

Put model calls behind one internal interface. Enforce allowed models, data handling rules, token limits, rate limits, retries, fallbacks, and per-tenant cost attribution there. Record model name, request ID, input and output token counts, latency, safety decisions, and errors.

Choose models by task rather than by brand: a smaller model may classify or extract fields, while a stronger model handles ambiguous planning. Define a fallback policy that preserves the output schema and does not silently reduce safety controls.

5. Tools are security boundaries

Every API, MCP server, database, SaaS connector, or code executor is a privileged capability. AWS identifies inbound and outbound authentication and authorization as separate concerns.

  • Use a workload identity for the agent and task-scoped credentials for tools.
  • Store secrets in a secret manager; never place them in prompts, logs, or source control.
  • Allow-list tool names and endpoints.
  • Validate arguments with strict schemas, ranges, ownership checks, and tenant filters.
  • Set timeouts, retry budgets, payload limits, and circuit breakers.
  • Make writes idempotent and require approval for irreversible or high-impact actions.
  • Validate tool results before passing them back to the model.
  • Log who authorized a call, which policy matched, and what happened.

For browser automation, isolate sessions and treat downloaded files, web content, and instructions returned by a site as untrusted input.

6. Knowledge retrieval and memory

Separate knowledge from memory. Retrieval answers questions from approved, indexed sources. Memory stores information about a session or user that the agent is permitted to retain.

Retrieval checklist

  • Ingest documents with source, owner, tenant, classification, and version metadata.
  • Filter by authorization before similarity search results reach the model.
  • Return citations or source IDs with every retrieved passage.
  • Re-index when source content changes and remove deleted material.
  • Measure recall and answer faithfulness with a fixed evaluation set.

Short-term and long-term memory

Keep the current turn and workflow state in session memory. Store durable preferences or facts only with an explicit retention policy, deletion path, and access check. Google recommends external persistent storage for production applications because in-memory state disappears when instances terminate.

CREATE TABLE agent_runs (
  run_id TEXT PRIMARY KEY,
  tenant_id TEXT NOT NULL,
  status TEXT NOT NULL,
  state JSONB NOT NULL,
  created_at TIMESTAMP NOT NULL,
  updated_at TIMESTAMP NOT NULL
);

CREATE INDEX agent_runs_tenant_status ON agent_runs (tenant_id, status);

7. Pick a runtime

Runtime Use it when Trade-off
Managed Agent Runtime You want built-in lifecycle, scaling, memory, identity, and observability Opinionated environment and less low-level control
Cloud Run You want containerized stateless services, custom tools, and scale-to-zero Attach external stores for state; manage more components yourself
GKE You need Kubernetes-level control, complex topology, or existing GKE operations Higher platform and operations overhead
Bedrock AgentCore AWS-native runtime, MCP gateway, memory, identity, observability, evaluations, and Cedar policy fit your stack More dependence on AWS services and controls

Google documents the managed runtime, Cloud Run, and GKE patterns in its architecture guide. Select based on control versus speed, portability, isolation, data residency, latency, and operating cost.

8. Deployment pattern

  1. Package the agent and tools as an immutable container.
  2. Inject configuration and secrets at runtime.
  3. Use a queue for long-running or retryable jobs; return a run ID to the caller.
  4. Persist checkpoints after retrieval, planning, and each tool call.
  5. Use separate identities and data stores for development, staging, and production.
  6. Release prompts, tool schemas, policies, and code together with a version ID.
  7. Roll out gradually and keep a rollback version.

Keep synchronous requests short. Streaming improves perceived latency, while asynchronous jobs protect the API from model or tool calls that exceed request timeouts.

9. Observability and evaluation

Normal CPU and memory metrics are not enough. AWS recommends tracing model calls, tool selections and results, failures, policy events, quality evaluations, latency, and cost.

Signal Record
Trace Run, model call, retrieval, policy decision, tool call, retry, handoff
Quality Task success, schema validity, groundedness, escalation rate, human rating
Reliability Timeouts, tool error rate, queue age, incomplete runs, retry count
Cost Tokens, model price, tool charges, storage, egress, cost per successful task
Safety Denied actions, prompt-injection detections, sensitive-data events, approvals

Build a regression set before changing prompts or models. Replay representative requests, compare structured outputs and policy events, and require human review for high-impact domains.

10. Performance, reliability, and cost controls

  • Reduce latency: stream responses, cache immutable retrieval results, parallelize independent reads, and avoid unnecessary model turns.
  • Bound work: cap steps, tool calls, tokens, wall-clock time, retries, and recursion depth.
  • Design for partial failure: use deadlines, exponential backoff with jitter, circuit breakers, and resumable checkpoints.
  • Control spend: route simple tasks to cheaper models, enforce per-user and per-run budgets, cache safely, and stop runs that exceed their plan.
  • Protect availability: queue bursts, apply backpressure, and keep a degraded response path when retrieval or a nonessential tool is unavailable.
  • Preserve correctness: prefer deterministic workflow steps for critical business logic and require confirmation before irreversible writes.

Agent requests can trigger several inference calls, tool invocations, memory lookups, and inter-agent communications; each adds latency, cost, and failure surface. See the AWS Agentic AI Lens.

11. Browser screenshots as an agent tool

If an agent needs visual evidence from a web page, define a narrow screenshot tool with an allow-listed URL policy, timeout, maximum image size, and retention rule. Capture only the page or element needed for the task, and treat page content as untrusted.

DIY browser tool considerations

  • Launch an isolated browser context per tenant or job.
  • Set a fixed viewport, device scale, locale, timezone, and user agent.
  • Wait for a selector or network idle instead of sleeping indefinitely.
  • Block unnecessary ads, trackers, and resource types to reduce cost and latency.
  • Hide sensitive selectors and avoid storing screenshots longer than required.
  • Classify bot checks, blank pages, timeouts, and failed loads as explicit outcomes.

12. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Use the API directly from an agent tool:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for the complete parameter list. It supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF paper size and page ranges, HTML/CSS to image, custom CSS and JavaScript, clicks, selector waits, delays, network idle, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Existing screenshot API parameter names also work, which simplifies migration.

ScreenshotNeo has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

13. Troubleshooting checklist

Symptom Likely cause Fix
Agent repeats a tool call No idempotency key or unclear result state Persist call status, use an idempotency key, and cap retries
Stale or unauthorized answers Retrieval runs before access filtering or indexing is old Filter by tenant and role before search; re-index changed sources
State disappears after deployment State stored only in process memory Persist checkpoints and memory in an external durable store
Requests time out Too many sequential model or tool calls Set deadlines, parallelize independent work, queue long jobs, and reduce context
Unexpected spend Unbounded loops, retries, or token budgets Set per-run budgets, maximum steps, model routing, and cost alerts
Unsafe action executes Tool authorization delegated to the model Enforce policy in code, validate arguments, and require human approval
Screenshot contains a popup Consent or widget cleanup was disabled or unsupported Enable cleanup, add a hide selector, or use custom CSS
Screenshot is blank or blocked Bot check, failed load, or page timeout Inspect verdict headers, adjust waits or headers, and treat the result as a non-success outcome

14. When to introduce multiple agents

Split into agents only when a boundary provides measurable value: distinct expertise, parallel work with independent failure handling, separate credentials or data domains, or an independently deployable team. Define the coordinator’s contract, message schema, timeout, retry policy, and authority for every handoff. Keep a single-agent fallback for routine requests.

15. Production readiness checklist

  • Agent charter is versioned and reviewed.
  • Critical logic is deterministic and replayable.
  • Every tool has an owner, schema, authorization policy, timeout, and audit event.
  • Secrets, identity, tenant isolation, and data retention are enforced outside the prompt.
  • Session state and checkpoints survive instance termination.
  • Model, tool, retrieval, safety, latency, and cost traces are searchable.
  • Regression evaluations run before releases.
  • Budgets, rate limits, backpressure, and escalation paths are configured.
  • Runbooks cover model outage, tool outage, stale data, unsafe output, and rollback.

FAQ

Should I use an agent framework?

Use one when it reduces orchestration code without hiding authorization, state, tracing, or failure behavior. A small explicit workflow is often easier to operate first.

Is a vector database required?

No. Use retrieval only when the agent needs changing or private source material. A relational store, search service, or document index can be sufficient.

Should every response be streamed?

Stream interactive text when it improves the user experience. Keep side effects and final status behind explicit, auditable workflow steps.

How do I choose between managed and self-managed hosting?

Choose managed hosting for faster delivery and built-in operations; choose containers or Kubernetes when portability, isolation, topology, or low-level control outweigh maintenance cost.

What is the first production metric?

Track successful task completion with latency and cost per successful task. Add safety, tool, retrieval, and quality dimensions immediately after.

The research framing in Infrastructure for AI Agents describes infrastructure as the systems and protocols that attribute actions, shape interactions, and detect or remedy harmful behavior. Those responsibilities should be visible in your architecture from the first release.