ScreenshotNeo

BlogEngineering

Build an AI Product Monitoring Tool

A practical guide to monitoring AI products with OpenTelemetry, quality evaluations, cost tracking, drift detection, and agent safety signals.

By the ScreenshotNeo team1 October 202611 min read

Direct answer: Build your monitoring tool around OpenTelemetry (OTel), then collect traces, metrics, and logs through an OTel Collector into a query backend. Instrument every model call, retrieval step, tool invocation, policy decision, and user-visible outcome. Track reliability, token cost, answer quality, behavior changes, safety, and business outcomes together. AI outputs are probabilistic, so latency and error dashboards alone cannot tell you whether the product is working.

Use a correlation ID for every user request and agent run. Record the model and provider, release and prompt versions, token counts, latency, retries, errors, retrieval sources, tool arguments and results, permissions, evaluator scores, and outcome identifiers. Apply redaction, retention, encryption, and access controls before storing prompts, responses, or tool payloads.

1. Use a portable monitoring architecture

OpenTelemetry is a vendor-neutral open-source framework for instrumenting, generating, collecting, and exporting traces, metrics, and logs. Its portable model lets you change storage or dashboard backends without rewriting application instrumentation. OTel reports support from more than 90 observability vendors. See the OpenTelemetry documentation.

Application or agent SDKs
          |
          v
   OpenTelemetry SDKs
          |
          v
   OpenTelemetry Collector
      /       |        \
     v        v         v
 traces    metrics     logs
      \       |        /
       v      v       v
  storage and query backend
          |
          v
 dashboards, evaluations, and alerts

Keep provider-specific fields in extensions, while retaining OTel-compatible core attributes. Store high-cardinality traces separately from roll-up metrics, and link every metric or evaluation score back to a trace or run ID.

Choose a backend

Decision axis Questions to answer
OTel compatibility Can it ingest standard traces, metrics, and logs without a proprietary SDK?
Cardinality and retention Can it query run, tenant, model, route, and release dimensions at your required retention period?
Evaluation support Can you store labeled examples, judge scores, experiments, and regression results?
Alerting Can alerts include representative traces and route to the people who can fix the issue?
Privacy and residency Where are prompts, outputs, and tool payloads stored, and how are they redacted?
Operations Does a hosted service reduce operational work, or does self-hosting provide needed control?

2. Define the event contract before writing instrumentation

Write a versioned event contract and privacy policy first. Decide which fields are retained, hashed, redacted, sampled, or excluded. Do not make raw prompts and tool payloads the only way to investigate a failure.

Field group Recommended fields Why it matters
Identity trace_id, run_id, request ID, tenant or cohort ID Joins one user request across model, retrieval, tools, and outcome events.
Deployment service, environment, release, route, prompt version, policy version Separates regressions caused by code, prompts, policies, or routing.
Model usage provider, model, input tokens, output tokens, estimated cost Explains cost changes and model-routing behavior.
Timing and reliability start time, duration, timeout, retry count, error type, queue time Supports latency percentiles and failure diagnosis.
Retrieval query ID, source IDs, ranking metadata, source changes Detects grounding and corpus regressions.
Tools tool name, arguments, permissions, result status, approval or denial Finds bad calls, loops, unauthorized actions, and slow tools.
Quality groundedness, relevance, completeness, schema validity, refusal correctness, tool-use correctness Measures whether a response is useful and safe.
Outcome user feedback, task success, conversion or workflow outcome ID Connects model behavior to product value.

Privacy checklist

  • Classify prompts, outputs, retrieved documents, and tool arguments before export.
  • Redact secrets, access tokens, payment data, and unnecessary personal information.
  • Hash identifiers when operators only need correlation, not identity.
  • Set retention by data class; keep aggregate metrics longer than raw content.
  • Encrypt data in transit and at rest, and restrict trace access by role and tenant.
  • Test deletion, export, residency, and audit requirements before production.

3. Instrument model, retrieval, tools, and outcomes

Create one parent span for the user request or agent run. Add child spans for model calls, retrieval, tool calls, post-processing, policy checks, and evaluation. Put stable dimensions on attributes and put large payloads in controlled events or linked storage.

user_request span
├── retrieval span
├── model_call span
├── tool_call span
├── policy_check span
├── post_process span
└── outcome span

Runnable Python example

The following example uses the OpenTelemetry SDK and console exporters so it runs locally without a backend. Replace the placeholder model and tool functions with your application calls, then configure an OTLP exporter for your Collector.

from time import perf_counter
from opentelemetry import trace, metrics
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, ConsoleSpanExporter
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import ConsoleMetricExporter, PeriodicExportingMetricReader

resource = Resource.create({
    "service.name": "ai-product",
    "service.version": "2026.10.01",
    "deployment.environment": "development",
})

tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(ConsoleSpanExporter()))
trace.set_tracer_provider(tracer_provider)
tracer = trace.get_tracer("ai-product.monitoring", "1.0.0")

reader = PeriodicExportingMetricReader(ConsoleMetricExporter(), export_interval_millis=5000)
metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))
meter = metrics.get_meter("ai-product.monitoring", "1.0.0")
requests = meter.create_counter("ai.requests")
failures = meter.create_counter("ai.failures")
latency = meter.create_histogram("ai.request.duration_ms")
input_tokens = meter.create_counter("ai.tokens.input")
output_tokens = meter.create_counter("ai.tokens.output")


def retrieve(question):
    with tracer.start_as_current_span("retrieval") as span:
        sources = [{"id": "doc-1", "title": "Example source"}]
        span.set_attribute("retrieval.source_count", len(sources))
        span.add_event("retrieval.completed", {"source_ids": "doc-1"})
        return sources


def call_model(question, sources):
    with tracer.start_as_current_span("model_call") as span:
        span.set_attribute("ai.provider", "example-provider")
        span.set_attribute("ai.model", "example-model")
        span.set_attribute("ai.prompt_version", "prompt-7")
        # Replace this with your provider SDK call.
        answer = "Replace this function with your model response."
        used_input_tokens = 120
        used_output_tokens = 24
        span.set_attribute("ai.input_tokens", used_input_tokens)
        span.set_attribute("ai.output_tokens", used_output_tokens)
        input_tokens.add(used_input_tokens, {"ai.model": "example-model"})
        output_tokens.add(used_output_tokens, {"ai.model": "example-model"})
        return answer, used_input_tokens, used_output_tokens


def handle_request(question, tenant="anonymous"):
    start = perf_counter()
    requests.add(1, {"route": "answer", "tenant": tenant})
    with tracer.start_as_current_span("user_request") as span:
        span.set_attribute("ai.route", "answer")
        span.set_attribute("ai.tenant", tenant)
        span.set_attribute("ai.release", "2026.10.01")
        try:
            sources = retrieve(question)
            answer, in_tokens, out_tokens = call_model(question, sources)
            with tracer.start_as_current_span("quality_evaluation") as evaluation:
                evaluation.set_attribute("evaluation.groundedness", 1.0)
                evaluation.set_attribute("evaluation.schema_valid", True)
            with tracer.start_as_current_span("outcome") as outcome:
                outcome.set_attribute("outcome.type", "answer_returned")
            return answer
        except Exception as exc:
            failures.add(1, {"route": "answer", "error.type": type(exc).__name__})
            span.record_exception(exc)
            span.set_attribute("error.type", type(exc).__name__)
            raise
        finally:
            latency.record((perf_counter() - start) * 1000, {"route": "answer"})


if __name__ == "__main__":
    print(handle_request("What changed in the latest release?", tenant="demo"))

Install the SDK packages with pip install opentelemetry-api opentelemetry-sdk. For production, export through an OTel Collector instead of writing directly from every application process.

cURL event example

curl -X POST http://localhost:4318/v1/traces \
  -H 'Content-Type: application/json' \
  -d '{"resourceSpans":[{"resource":{"attributes":[{"key":"service.name","value":{"stringValue":"ai-product"}}]},"scopeSpans":[{"spans":[{"traceId":"5b8aa5a2d2c872e8321cf37308d69df2","spanId":"051581bf3c55c1a1","name":"model_call","kind":1,"startTimeUnixNano":"1720000000000000000","endTimeUnixNano":"1720000000200000000","attributes":[{"key":"ai.model","value":{"stringValue":"example-model"}},{"key":"ai.input_tokens","value":{"intValue":"120"}},{"key":"ai.output_tokens","value":{"intValue":"24"}}]}]}]}]}'

Node.js span example

import { trace } from '@opentelemetry/api';

const tracer = trace.getTracer('ai-product');

export async function monitoredModelCall(question) {
  return tracer.startActiveSpan('model_call', async (span) => {
    const started = Date.now();
    span.setAttribute('ai.provider', 'example-provider');
    span.setAttribute('ai.model', 'example-model');
    span.setAttribute('ai.prompt_version', 'prompt-7');
    try {
      // Replace with your provider SDK call.
      const result = { answer: 'model response', inputTokens: 120, outputTokens: 24 };
      span.setAttribute('ai.input_tokens', result.inputTokens);
      span.setAttribute('ai.output_tokens', result.outputTokens);
      return result;
    } catch (error) {
      span.recordException(error);
      span.setAttribute('error.type', error.constructor?.name || 'Error');
      throw error;
    } finally {
      span.setAttribute('ai.duration_ms', Date.now() - started);
      span.end();
    }
  });
}

4. Build the five monitoring layers

Reliability

Track request volume, error rate, timeout rate, retry rate, queue depth, and end-to-end latency percentiles. Break each metric down by route, model, provider, release, tenant, and region where those dimensions are useful. Always show the denominator and time window.

Cost

Record input and output tokens on every model span. Calculate estimated cost using a versioned price table, then aggregate by feature, tenant, route, model, and release. Alert on sustained cost per successful task, not only total spend; traffic changes can make total spend rise while efficiency improves.

Quality

Score groundedness, relevance, answer completeness, schema validity, refusal correctness, and tool-use correctness. Use labeled examples or judge models, retain the evaluator version, and link every score to the original run. Keep an offline regression suite for releases and run continuous or sampled evaluations after deployment.

Behavior

Monitor retrieval-source changes, tool-call loops, unexpected permissions, fallback frequency, and input or output distribution shifts. Establish baselines by model, route, tenant, and release. Alert on sustained deviations and attach representative traces to each alert.

Safety and governance

Capture policy decisions, prompt-injection indicators, data-exfiltration signals, sensitive-content handling, and human approvals. A successful HTTP response can still be a safety failure, so safety outcomes need their own metrics and alerts.

5. Dashboards and alerts that lead to action

Dashboard Minimum panels Drill-down
Reliability Throughput, error rate, timeout rate, p50/p95/p99 latency, retries, queue depth Trace by release, model, route, and provider
Cost Tokens, estimated spend, cost per request, cost per successful task Prompt version, tenant, model routing, and long-running traces
Quality Groundedness, relevance, completeness, schema validity, refusal correctness Evaluation examples and source documents
Agent behavior Tool success, loops, fallbacks, permission denials, retrieval changes Tool arguments, approvals, and child spans
Safety Injection indicators, sensitive-content events, policy decisions, human approvals Redacted trace and review history

Define alert conditions with a baseline, a minimum sample size, and a sustained window. For each alert, include the changed dimension, a comparison period, and representative trace IDs. Test alert delivery and escalation, not just the query.

6. Regression tests and release gates

  1. Keep a labeled set covering normal requests, difficult questions, refusals, tool use, retrieval failures, and adversarial inputs.
  2. Run the set against the candidate release and record model, prompt, policy, evaluator, and dataset versions.
  3. Gate deployment on agreed quality and safety thresholds plus reliability and cost budgets.
  4. Sample production runs for continuous evaluation after release.
  5. Compare cohorts by model, route, tenant, and release so an average score cannot hide a regression.

7. OpenSearch as one implementation option

The OpenSearch GenAI observability guide describes a path using Python SDK instrumentation, an OTel Collector, local evaluation, middleware processing, OpenSearch dashboards, trace inspection, and quality scoring. Its SDK documents functions including register(), @observe, enrich(), score(), and evaluate(), with automatic tracing for OpenAI, Anthropic, Bedrock, LangChain, and more than 20 libraries. The guide lists Python 3.10+ and Docker prerequisites. It reports that traces typically appear 2–5 seconds after the batch processor flushes. Treat these implementation details as release-sensitive; verify them in the current OpenSearch documentation.

8. Reliability, performance, and cost notes

  • Sampling: retain complete traces for errors, safety events, and evaluation failures; sample routine successful traces according to your investigation needs.
  • Payload size: avoid putting large prompts, outputs, and documents in every span attribute. Store redacted content separately and link it by ID.
  • Exporter failures: use bounded queues, retries, and a drop policy so telemetry cannot block user requests indefinitely.
  • Cardinality: keep unbounded values such as raw user text out of metric labels. Put them in traces or logs with access controls.
  • Latency: batch export asynchronously and measure instrumentation overhead in your own workload.
  • Cost: estimate both model spend and observability spend. Retention, high-cardinality queries, evaluator-model calls, and raw payload storage can dominate.
  • Reliability: test missing telemetry, Collector outages, schema changes, alert delivery, and retention jobs as failure modes.

9. Troubleshooting

Symptom Likely cause Fix
No traces arrive Exporter endpoint, protocol, or credentials are wrong. Send a known-good test span, inspect Collector logs, and verify OTLP HTTP versus gRPC configuration.
Traces are split across runs Trace or correlation context is not propagated through async jobs or tools. Propagate context explicitly and preserve the run ID in every child operation.
Metrics are expensive or unusable High-cardinality values are metric labels. Move raw IDs and text to traces; keep metric dimensions bounded.
Quality scores cannot be explained The evaluator, dataset, or prompt version was not recorded. Store evaluator version, rubric, example ID, and linked trace with every score.
Costs rise unexpectedly Longer prompts, retries, model fallback, or routing changes. Compare token counts and retry/fallback spans by release, route, and model.
Agent loops go unnoticed Tool calls are logged only as application messages. Create a span per tool call and alert on repeated calls, duration, and permission changes.
Privacy review blocks rollout Raw prompts or outputs are exported without controls. Redact, hash, restrict, encrypt, and set retention before enabling payload capture.
Alerts are noisy Single events trigger alerts without a baseline or sample threshold. Alert on sustained deviation and include cohort filters and representative traces.

10. Add visual monitoring for AI product interfaces

If your AI product has a web interface, visual checks can catch broken streaming states, missing citations, empty result panels, or a consent popup covering the answer. Run them alongside telemetry and associate each capture with a release, route, and test case. A screenshot should support an alert; it should not replace traces and evaluations.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server for developer workflows. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

Use the ScreenshotNeo documentation for the complete option list, including full-page capture with lazy images loaded, CSS element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, async jobs, webhooks, bulk capture, usage, and the OpenAPI specification.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Do I need a vendor-specific AI monitoring SDK?

No. Use OTel as the collection layer and keep provider-specific attributes as extensions. This preserves portability while retaining model and agent context.

Should every prompt and response be stored?

No. Define retention and redaction rules first. Keep aggregate metrics longer, and retain raw content only when it is necessary for investigation and permitted by your policy.

How do I detect model drift?

Compare input, output, retrieval, tool, quality, and cost distributions by model, route, tenant, and release. Establish baselines and alert on sustained deviations.

What makes an AI alert actionable?

It should identify the changed cohort and time window, show the denominator, state the threshold, and link to representative traces and evaluation examples.

Can screenshots replace AI evaluations?

No. Screenshots verify the rendered interface. Traces, structured events, and evaluations explain model quality, tool behavior, safety, and cost.