Functional AI is coming soon. Join the waitlist for early access.

// blog

8 min read

What Is an Eval?

An eval is an automated test that measures whether your AI prompt or agent does what you intend—dataset, task, scorer—and how to gate CI/CD on the result.

An eval (short for evaluation) is a structured, repeatable test that measures whether your AI prompt or agent's outputs meet a defined quality standard. You give it a dataset of inputs, specify what the prompt or agent should do with each one, and score the outputs against a threshold — automatically, on every deploy. If the score drops below the threshold, the build stops.

The three components of every eval

Every eval has the same structure regardless of how complex the prompt or agent is.

Dataset — a versioned collection of test cases. Each case has an input and, optionally, an expected output. Fifty well-chosen examples drawn from real user queries will catch large regressions. Two hundred examples gives statistical confidence on quality differences as small as 3–5%. Synthetic examples generated from documentation look clean but miss the edge cases that real users find — the most effective datasets combine human-crafted edge cases, real production samples with PII removed, and synthetic expansions for underrepresented scenarios.

Task — the prompt or agent itself. In an eval, the task is whatever your prompt or agent does: classify a support ticket, extract entities from a contract, generate a draft reply. The eval calls it with every example in the dataset and collects the outputs.

Scorer — the logic that decides whether each output is good. The scorer takes the input, the output, and optionally the expected output, then returns a score. This is where evals get interesting, because there are three distinct approaches — and each one has a failure mode.

Why evals catch what unit tests miss

Unit tests check deterministic correctness: either parseDate("July 4th") returns the right timestamp or it doesn't. Prompts and agents don't behave that way. The same prompt, the same model, the same input can produce subtly different outputs across runs. And the definition of "correct" is often not a binary condition — it's a threshold on a distribution.

More importantly, there are two distinct ways a prompt or agent can break:

  1. You change something. You edit a prompt, swap a model version, adjust a retrieval config.
  2. Something changes you. A model provider silently updates weights without changing the API endpoint. GPT-4 Turbo's underlying model weights changed multiple times in 2023 and 2024 without the API version identifier changing. Anthropic's August 2025 postmortem documented routing errors affecting up to 16% of Claude Sonnet 4 requests — with no API modification.

Unit tests catch the first category on a good day. They catch nothing in the second. Evals catch both — because the scorer runs against real output, not a mocked return value.

Offline evals vs. online evals: two different problems

These are not the same thing, and conflating them is how teams end up with gaps in coverage.

Offline evals run before deployment, against a fixed versioned dataset. You hold everything constant — model version, prompt hash, retrieval config — and measure the prompt or agent's behavior on known inputs. This is your CI quality gate. If an offline eval fails, the deploy stops.

Online evals run in production. You sample 5–10% of real traffic, score it with an automated evaluator, and monitor for drift over time. This is the only layer that catches changes that happen to you: model provider updates without a version increment, input distribution shifts as new user segments arrive, behavior patterns you didn't anticipate when you built your test set.

Both are necessary. Offline catches regressions you introduce. Online catches regressions that arrive uninvited. Teams that run only offline evals eventually discover a production incident with no way to determine when it started or what caused it.

Three scoring approaches — and where each one fails you

Deterministic scoring

Exact match, JSON schema validation, regex checks. Fast, free, fully predictable. Use these first — they run in milliseconds and cost nothing per run.

function exactMatch(output: string, expected: string): number {
  return output.trim().toLowerCase() === expected.trim().toLowerCase() ? 1 : 0;
}

Where it fails: it can't judge whether a support reply is actually helpful, or whether a generated summary stays faithful to the source document. For anything subjective, you need a different approach.

LLM-as-judge

A second model scores the output against a rubric. Flexible for open-ended quality. Where it fails: the judge is itself a model, so its scores vary across runs. An eval with a faithfulness judge that scores 0.82 one run and 0.79 the next isn't broken — it's probabilistic. Track trends rather than point scores, and validate the judge against human-reviewed examples before trusting it to gate a deploy.

Human review

The most accurate approach, and the slowest. Use it to calibrate your automated scorers, not as a primary CI gate. Score a reviewed sample of 20–50 outputs, compare your automated scorer results against human labels, and set thresholds that match human judgment. Then automate everything forward.

In practice, most prompts and agents need all three: deterministic scorers catch the schema breaks, LLM-as-judge covers subjective quality, and periodic human review validates that the automated scorers are still calibrated.

The failure mode nobody builds for: category-level regression

Iterative prompt improvement is a silent accumulator. You fix a hallucination on refund requests → overall pass rate climbs. You tighten tone instructions → score goes up again. You ship v4 of the prompt, and the negation accuracy on a specific cancellation query category collapses from 0.94 to 0.31 — invisible inside the aggregate pass rate.

This is not an edge case. It's the standard outcome of improving prompts without category-level tracking. The fix is to score by category, not just by total. If your prompt or agent handles three distinct query types, it needs three distinct scorer tracks. An aggregate score tells you whether the prompt or agent improved on average. Category scores tell you which user segment you just broke.

Wiring an eval into CI/CD

The pattern is straightforward: scope the trigger to files that affect prompt or agent behavior, run the eval, exit non-zero on threshold failure.

# .github/workflows/eval.yml
on:
  push:
    paths:
      - 'prompts/**'
      - 'fn.yml'
      - 'src/functions/**'
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci
      - run: fn eval classify@v3 --threshold 0.95

With Functional AI, fn eval classify@v3 calls your prompt or agent against every example in your eval dataset, scores each output, and exits non-zero if any scorer falls below its threshold:

Running eval: classify@v3
Dataset:       142 examples
Scorers:       exact_match, llm_judge_tone
──────────────────────────────────────────
exact_match      0.98   ✓  pass (threshold 0.95)
llm_judge_tone   0.91   ✗  fail (threshold 0.95)
──────────────────────────────────────────
Eval failed. 3 examples below threshold:
  ex-0042   support/refund         score 0.71
  ex-0187   support/cancellation   score 0.68
  ex-0291   support/billing        score 0.74

Three failing examples with a category and a score. That's a fixable problem, not a mystery incident. The PR doesn't merge until the eval passes.

What a passing eval actually means

A passing score is not a guarantee — it's a contract. You are asserting that on the test cases you've curated, at the threshold you've chosen, this prompt or agent's behavior meets your definition of good. The eval is only as strong as the dataset and the scorer behind it.

This is also how the ownership problem gets solved. When a prompt or agent's eval lives alongside its code — versioned dataset, explicit scorer thresholds, CI gate — there's no ambiguity about who is responsible for quality. The team that ships the prompt or agent owns the eval. That contract is checked on every push, not assumed in a Notion doc.

In Functional AI, evals are part of the deployment contract. fn eval classify@v3 runs the scorer suite before fn deploy promotes the version. The runtime enforces the threshold; your team defines what good means. That separation — runtime enforces, team defines — is what makes AI features trustworthy in production.

Join the waitlist to wire evals into your deployment pipeline.

FAQ

What is an eval in AI?

An eval (evaluation) is an automated test that measures whether an AI prompt or agent's outputs meet a defined quality standard. It runs a dataset of inputs through the prompt or agent, scores each output with one or more scorers, and fails the build if any score falls below its threshold.

What is the difference between an eval and a unit test?

Unit tests check deterministic correctness — a given input produces exactly one correct output. Evals handle nondeterminism and subjective quality. They measure whether outputs are good enough across a distribution, using scorers that can be deterministic (exact match, schema validation), LLM-based (judge model with a rubric), or human-reviewed.

How many examples do I need in an eval dataset?

Fifty well-chosen examples drawn from real user queries will catch large regressions. Two hundred gives statistical confidence on differences as small as 3–5%. More than 500 shows diminishing returns unless your prompt or agent has highly varied subtasks that need separate coverage.

What is LLM-as-judge scoring?

LLM-as-judge uses a second language model to score the output of your AI prompt or agent against a rubric. It's useful for open-ended quality that can't be measured deterministically. Because the judge is itself nondeterministic, track score trends across runs rather than treating a single result as a hard pass or fail.

What is the difference between offline and online evals?

Offline evals run before deployment on a fixed dataset — they are your CI quality gate, catching regressions you introduce. Online evals run in production on sampled real traffic — they catch model provider drift and input distribution shifts that a fixed dataset cannot see. Both are necessary.

Who should own the eval?

The team that ships the prompt or agent owns the eval. That means maintaining the dataset, setting threshold values, and keeping the CI gate active. When the eval lives in version control alongside the prompt or agent code, ownership is explicit rather than assumed.

// private beta

Ship agents you can trust.

Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.

No spam. One email when Functional AI launches.

Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.

What Is an Eval? — Functional AI