Functional AI is coming soon. Join the waitlist for early access.

// blog

6 min read

AI Observability: A Production Engineer's Guide

AI observability captures traces, eval scores, and cost signals from agents at runtime—catching failures that HTTP 200 hides. A concrete engineering guide.

AI observability is the practice of capturing trace, quality, and cost signal from your agents at runtime so you can answer "why did my agent do that?" for any production request — including the ones that returned HTTP 200 but were wrong. It is not a superset of APM; it is a different discipline, because the failure modes are semantic, not operational.

Why AI observability is different from traditional APM

A conventional service fails in ways your monitoring stack sees: exceptions, timeout codes, high error rates, latency spikes. An agent can fail in ways your monitoring stack is completely blind to — it returns HTTP 200 at normal latency with a confidently wrong answer.

Research consistently finds that inaccuracy is the generative AI risk organizations most often experience in practice — ahead of security or privacy incidents. Most production monitoring stacks are built to catch the latter and are blind to the former.

The math makes this urgent. An agent running at 85% per-step accuracy — which sounds high — succeeds only about 20% of the time on a 10-step workflow: 0.85¹⁰ ≈ 0.20. Each step's error probability multiplies with the next. Most teams don't internalize this until a multi-step workflow fails in a way that looks random but is structurally inevitable.

And silent failures stay silent. Production LLM monitoring research finds that the majority of silent quality failures surface only when a person reads the output — not through any automated test, health check, or alert. That lag runs to days or weeks, and by then users have already formed an opinion.

The four failure modes AI observability must catch

Silent quality degradation. Output quality drifts downward across days or weeks with no triggering incident and no error in the logs. A prompt change that looked safe in development silently degrades a user-facing flow. Token spend can double from one prompt edit that nobody flagged as meaningful. Quality alerts need a rolling window on sampled production traffic — a single-point alert on aggregate score misses gradual drift entirely.

Prompt regression. A new prompt version ships. Quality on the classification task drops from 94% to 79%. You find out three days later from a support ticket. The fix is a CI/CD quality gate that scores the candidate prompt against a golden eval set before it merges.

Tool cascade failure. In a multi-step workflow, one tool call returns malformed output. The next LLM call consumes it as context and propagates the error. Each step's error compounds: what looks like a bizarre final answer started as a bad tool schema two steps back. You need span-level visibility into every tool call — input, output, latency, success/fail status — not just the final response.

Cost explosion. Development runs 100 requests at $0.02 each: negligible. Production runs 50,000 requests per month with longer context windows, at $0.18 each: $9,000 per month that was never in the business case. Without per-request cost tracking broken down by token type, you get a surprise bill before you get a dashboard.

The signal stack: what to instrument

The OpenTelemetry GenAI Semantic Conventions, developed by the OTel GenAI SIG since April 2024, define a standard gen_ai.* attribute vocabulary for AI workloads. Using them means your instrumentation works with any compatible backend without custom parsing rules, and a span from a LangChain agent looks identical to one from a raw OpenAI call.

The span hierarchy for a multi-agent Agentic Function looks like this:

invoke_agent classify-document  (INTERNAL)
  chat gpt-4o                   (CLIENT)
    gen_ai.usage.input_tokens:  1842
    gen_ai.usage.output_tokens: 67
    gen_ai.client.operation.duration: 834ms
  execute_tool fetch-policy     (INTERNAL)
    duration_ms: 234
    tool_call.success: true
  chat gpt-4o                   (CLIENT)
    gen_ai.usage.output_tokens: 412
    gen_ai.response.finish_reasons: ["stop"]
  execute_tool write-result     (INTERNAL)
    tool_call.success: true

The attributes that matter for production debugging:

AttributeWhy it matters
gen_ai.request.modelWhich model you asked for
gen_ai.response.modelWhich model actually ran (differs on fallback)
gen_ai.usage.input_tokensPrompt token cost driver
gen_ai.usage.output_tokensCompletion token cost driver
gen_ai.usage.reasoning_tokensReasoning cost (o-series models)
gen_ai.client.operation.durationPer-span latency for bottleneck identification
gen_ai.response.finish_reasonsWhy the model stopped: stop, tool_calls, length

One attribute most teams skip: cached input tokens. GPT-4o and Claude 3.5 Sonnet both support prompt caching for repeated context prefixes. Tracking cached vs. uncached tokens separately is the difference between understanding your real cost per request and guessing at it — the savings are large enough to justify the instrumentation work on their own.

One rule on content: store prompt and completion text as span events, not span attributes. Events can be filtered or dropped at the OTel Collector level for PII compliance; attributes are always indexed and always exported. A 4,000-token system prompt stored as a span attribute floods your backend and creates a compliance surface.

Evals are observability too

Traces tell you what happened. Evals tell you whether what happened was any good. Both are required; neither replaces the other.

Offline eval runs in CI before the deploy ships. In Functional AI, that looks like:

fn eval classify@v3 --dataset golden-set.jsonl --threshold 0.95

This runs your Agentic Function against a versioned eval set and blocks the deploy if the pass rate drops below 95%. It is the equivalent of a unit test for output quality. If your eval set is representative of real traffic, this gate catches prompt regressions before they reach users — join the waitlist at /#waitlist to see this integrated into the Agentic Function Runtime.

Online eval runs on sampled production traffic after the deploy. Scoring 5% of live requests against a task-specific rubric, trended over a 7–14 day rolling window, gives you a quality trend line that detects directional drift before it becomes user-visible. A growing cluster of outputs near your confidence floor often precedes a measurable accuracy drop by several days. That lead time is the entire point.

The eval-to-observability loop:

  1. Trace every Agentic Function invocation — spans, tool calls, model choices, token counts
  2. Score a sample of production traces with your eval rubric
  3. Alert on sustained directional drift over a rolling window, not single-point variance
  4. Gate the next version in CI against the scoring baseline before it merges
  5. Ship knowing the new version does not regress what the old one did well

Model fallback: what to observe

When your primary model is unavailable and your runtime falls back to a secondary — or when a model provider silently updates behavior between API versions — your traces must capture which model actually ran on each span, not just which model you requested.

This is where gen_ai.request.model and gen_ai.response.model matter as distinct attributes. If they differ on a span, a fallback occurred. If they are always identical in your traces, you are not tracking fallback events and cannot correlate quality changes with model substitutions.

In Functional AI, the Agentic Function Runtime handles model fallback automatically and surfaces both the requested and actual model on every span. When eval scores dip after a fallback event, you can see it in the trace, then decide whether to pin the model, tune the prompt for the fallback model, or adjust your fallback threshold.

The three quality dimensions every AI dashboard needs

Quality: Eval scores trended over time, segmented by prompt version and model version. This is the dimension standard APM is blind to, and where silent failure hides. Watch the score distribution, not just the average — a growing cluster of outputs near your confidence floor precedes an accuracy drop by days.

Cost: Per-request spend broken down by prompt tokens, completion tokens, cached tokens, and reasoning tokens. Identify which workflows or users consume a disproportionate share of token budget. Track cost per agent version so you can compare the quality-to-cost ratio of a new model against the old one before committing to it.

Latency: Time to first token and total request duration per span, so you can identify whether a bottleneck is the LLM call, a slow tool, or a retrieval step. Users form a trust judgment in the first two seconds of a response, before they have read a word.

A monitoring stack that covers only latency — because it looks like classic APM — misses the failures that actually kill AI products in production.

FAQ

What is AI observability?

AI observability is the practice of capturing trace, quality, and cost signals from AI agents at runtime. It answers "why did my agent produce that output?" for any production request, including requests that returned HTTP 200 but were semantically wrong. It is distinct from traditional APM, which monitors operational health but cannot see semantic failures.

How is AI observability different from traditional application monitoring?

Traditional monitoring catches operational failures: exceptions, timeouts, error rates, latency spikes. AI observability catches semantic failures: hallucinations, wrong answers, goal drift, quality regressions. An agent can return HTTP 200 at normal latency with a confidently wrong answer — no APM tool flags that without semantic evaluation.

What is a trace in AI observability?

A trace is the complete execution record of one agent invocation. It contains nested spans — one per LLM call, tool call, or retrieval step — each carrying attributes like model name, token counts, latency, and inputs/outputs. Traces let you reconstruct exactly what the agent did, in what order, and where it went wrong.

What should I instrument first for AI observability?

Start with three signals: model name and token counts on every LLM call, input/output and success status on every tool call, and an end-to-end latency span wrapping the full agent invocation. Use OpenTelemetry gen_ai.* attribute names from the start so your traces are compatible with any backend. Add eval scores next.

How do evals fit into AI observability?

Evals score output quality systematically, turning subjective quality judgments into trackable metrics. Run them in CI as quality gates before deploy, and on a sampled fraction of production traffic to detect drift over rolling time windows. Without evals, observability only shows you what happened — not whether it was right.

What is the OpenTelemetry GenAI semantic convention?

It is the standard attribute vocabulary for LLM and agent instrumentation, maintained by the OTel GenAI SIG since April 2024. Using gen_ai.* attribute names ensures traces are compatible with any OTel-compliant backend without custom parsing. Key attributes include gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and gen_ai.client.operation.duration.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.