Functional AI is coming soon. Join the waitlist for early access.

// blog

6 min read

Eval Metrics: What to Measure, What to Gate, What to Skip

Which eval metrics to run in CI, how to calibrate thresholds, and the agentic failure modes most guides skip.

Eval metrics are the scored signals that convert subjective agent quality into a number your CI pipeline can act on. Each metric takes an input, an output, and optionally an expected value, and returns a score between 0 and 1. A suite of metrics is what separates "the prompt feels better" from "this Agentic Function passed 94 of 100 test cases at ≥0.90 factuality before it was allowed to merge." Without them, every regression surfaces as a user complaint instead of a failing check.

What are the three types of eval metrics?

Metrics fall into three families ordered by stability and cost. Use the cheapest type that answers the question.

Deterministic metrics run in microseconds, cost nothing, and return the same answer every time. They cover everything with a single correct answer:

  • Schema / JSON validity — does the output parse? Gate at 1.0; any failure blocks the deploy.
  • Exact match — does the extracted entity match the expected string?
  • Regex / pattern — does the output contain a required field or format?
  • Length constraint — is the response within the character budget?

Run these on 100% of test cases on every commit. They catch obvious regressions for free.

Embedding-based metrics convert output and reference into vectors and measure cosine distance. A semantic similarity score ≥0.85 typically means the response conveys the same meaning even if the phrasing differs. Cost is roughly $0.0001 per call. Use them when exact match is too brittle — multi-path answers, paraphrase outputs, or any case where several wordings are equally valid.

LLM-as-judge metrics use a model to score qualities that code can't measure: factuality, instruction-following, tone, reasoning quality. They're the most expressive and the most variable. A judge running at temperature 0.0 still produces different scores on identical inputs 5–15% of the time in practice. Three mitigations:

  1. Write a binary rubric, not an open-ended question. "Rate from 1–5" drifts; "Return 1 if every claim is supported by the provided context, else 0" does not.
  2. Run the judge at temperature 0. Even a small positive temperature introduces compounding variance across a 200-case suite.
  3. Calibrate against human labels before using as a gate. Score 100 cases manually. If your judge agrees with humans on fewer than 85% of them, fix the rubric first.

Our post on LLM-as-judge covers scorer design in depth.

What eval metrics do agents need that LLM guides skip?

Most eval metric guides are written for single-turn calls — prompt in, response out. Agents are different. An agent executes a trajectory: planning, tool calls, retries, state accumulation. The final answer is only the last artifact in a longer chain, and each step is a new failure surface.

Tool selection accuracy measures whether the agent called the right tools. Score = expected tools called / expected tools total. Useful — but fatally insufficient on its own. An agent can pick the right function from a registry of 28 and still fail the task completely: tool-selection accuracy logs 1.0 while task completion logs 0. What broke? Could be the arguments, the response handling, the retry loop — tool selection can't tell you.

Argument correctness is the layer below. Did the agent pass the right arguments? A travel agent that calls search_flights with departure_date="next Friday" instead of "2026-09-26" will get a 400 on every attempt. Tool selection: 1.0. Booking: never happened. Argument correctness catches this; tool selection doesn't.

Trajectory fidelity scores the full step sequence — not just whether the right tools appeared, but whether they appeared in the right order with valid arguments and without policy-violating intermediate actions. A correct final answer reached in 20 steps with two compliance violations in the middle is a failing trajectory, even if final-answer eval gives it a 1.0.

Step efficiency (actual steps / minimum steps) flags agents that complete the task through an unnecessarily long path. An agent that takes 12 steps where 4 would do is expensive today and fragile tomorrow: each extra step is another place the next model update can introduce a regression.

pass^k measures reliability across k independent rollouts of the same task. Run a hard case 8 times; the fraction succeeding on all 8 is your pass^8 score. Even top models land below 25% at pass^8 on complex multi-turn tasks. Don't run pass^k on every PR — 30 cases × 8 rollouts is 240 agent runs. Reserve it for release candidates, where it catches planner regressions before they ship.

How do you calibrate CI thresholds?

Thresholds should come from your data, not a heuristic. Run your eval suite on the last known-good version of your Agentic Function, observe the score distribution, and set the gate 2–3 points below the p10 score. That gives you a regression signal that fires on real breaks, not natural variance.

Starting points if you're building from zero:

MetricGateNotes
JSON / schema validity1.0Zero tolerance — parse failures block
Exact match1.0For classification and extraction tasks
Semantic similarity≥0.85Paraphrase-safe regression baseline
Factuality (LLM-judge)≥0.85Below 0.80, users notice hallucinations
Tool selection accuracy≥0.90Always pair with argument correctness
Task completion≥0.80End-to-end pass/fail on trajectories

In Functional AI you declare these thresholds in fn.yml and the runtime enforces them on every deploy — the previous version keeps serving production traffic until the new one passes:

evals:
  suite: support-agent@v4
  thresholds:
    schema_validity: 1.0
    factuality: 0.85
    tool_selection: 0.90
    task_completion: 0.80

Run the suite locally before pushing:

fn eval support-agent@v4 --threshold 0.85

If any metric misses its gate, the deploy is blocked and the previous version keeps serving. Sign up for early access to test this behavior against your own Agentic Function.

What failure modes do standard guides skip?

The compound average trap. Averaging factuality (0.95) and JSON validity (0.0) gives 0.475 — which looks like a recoverable regression, not a total production breakage. Gate on the minimum across all metrics, not the mean. A single 0.0 on a binary check is a hard stop, not something to offset with a high score elsewhere.

Metric flakiness. A judge that returns 0.85 for input A on one run and 0.72 on the next generates noise, not signal. Track the standard deviation of your judge across 20 identical inputs. If it exceeds 0.05, the rubric needs tightening before that metric earns gate status in CI.

Scorer drift. Changing the judge prompt, upgrading to a newer judge model, or rewording a rubric shifts the score distribution. A factuality score of 0.82 from judge-v1 is not comparable to 0.82 from judge-v2. Version your scorers. When you update one, re-score a calibration slice of 100 cases to establish the new baseline before enforcing the threshold — otherwise you'll flag a non-existent regression or miss a real one.

FAQ

What's the first eval metric I should add? Schema validity — a deterministic check that your agent's output parses correctly. It costs nothing to run, catches format regressions immediately, and gates at 1.0 with no threshold calibration needed.

How many eval metrics should I track? Start with three: one deterministic (schema validity or exact match), one embedding-based (semantic similarity), and one LLM-judge (factuality or task completion). Add more only when you have a production failure that those three can't diagnose.

Should I use the same model as my production agent to judge outputs? No. The judge and the agent should be different models. A model judging its own outputs shows self-preference bias — it rates its own generations more favorably than independent human raters do on the same examples.

What threshold should I set for factuality? Run your eval suite on the last known-good version, observe the p10 score, and set the gate 2–3 points below it. A common starting point is ≥0.85 — below 0.80, users reliably notice hallucinations in production.

How do I stop my LLM-as-judge from being flaky? Run the judge at temperature 0, write the rubric as a binary condition rather than a subjective scale, and validate agreement against 100 human-labeled cases before using it as a CI gate.

What's the difference between tool selection accuracy and task completion? Tool selection accuracy measures whether the right tools were called; task completion measures whether the agent actually achieved the goal. An agent can score 1.0 on tool selection and 0.0 on task completion if the arguments were wrong, the model ignored the tool response, or the retry loop ran without recovering. You need both.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.