Functional AI is coming soon. Join the waitlist for early access.

Evaluating an LLM prompt before shipping

Evaluating an LLM prompt before shipping means running it against a labelled dataset and checking whether the output meets a defined quality threshold — so regressions are caught before users see them. Functional AI turns that check into a single CLI command: set a dataset and a threshold, and every version of your prompt, agent, or workflow is scored before it reaches production.

Evaluation in Functional AI is a command, not a notebook. You define a dataset and a threshold; Functional AI turns every change into a pass/fail signal.

Dataset format

A dataset is a JSONL file — one JSON object per line — with an input and an expected field:

{"input": "ticket-001.pdf", "expected": {"category": "billing", "priority": "high"}}
{"input": "ticket-002.pdf", "expected": {"category": "support", "priority": "low"}}

expected may be a full output object or a partial one. Only the keys you specify are scored, so you can assert on category alone while ignoring confidence.

Scoring

fn eval classify --dataset tickets.jsonl --threshold 0.95

Each case is graded field-by-field. The reported score is the fraction of asserted fields that matched. A score below --threshold exits non-zero.

Open-ended output

Field matching works when the output is structured. For free-text or agent output, exact match is the wrong tool: use model-graded checks instead. Treat those scores as a signal, not ground truth — LLM judges carry positional and self-preference bias and drift between runs. Functional AI tracks graded scores per version so drift is visible, but a human review cadence stays the calibration layer that keeps automated scoring honest.

Multi-run consistency

Models are non-deterministic. Use --runs to execute each case multiple times and measure variance:

fn eval classify --dataset tickets.jsonl --runs 5

The report includes a consistency figure — how many of the N runs produced identical output — alongside the score.

Regression detection

Every eval is compared to the last recorded run for that Agentic Function. The report shows the delta versus the previous version, so a quiet 2% regression doesn't slip through.

Running LLM evals in CI

Because fn eval exits non-zero below threshold, gating a deploy is one step:

- name: Evaluate function
  run: fn eval classify --dataset tickets.jsonl --threshold 0.95

If the eval fails, the job fails, and the deploy never runs. No custom failure logic, no post-deploy rollback — the bad version never reaches users.

Testing an AI agent before releasing it to users

An agent is harder to evaluate than a single prompt: it calls tools, spans multiple turns, and its output may vary across runs. The evaluation approach is the same — define what success looks like, run it against test cases, gate on a threshold — but consistency matters more.

Write test cases as a JSONL dataset where input describes the starting state and expected captures the required outcome:

{"input": {"user_message": "cancel my subscription"}, "expected": {"action": "cancel", "confirmed": true}}
fn eval classify --dataset agent-cases.jsonl --threshold 0.90 --runs 3

Check both the score and the consistency figure. An agent that returns the correct answer two runs out of three is not production-ready. Gate the deploy when both metrics clear the threshold.

Functional AI tracks the score delta against the previous version, so a quiet regression between agent releases doesn't slip through undetected. To evaluate your own agents before they reach users, join the waitlist.

FAQ

How do I evaluate an LLM prompt before shipping it?

Evaluating an LLM prompt before shipping means running it against a labelled dataset and checking whether the output meets a defined quality threshold — so regressions are caught before users see them. In Functional AI: create a JSONL file of input/expected pairs and run fn eval classify --dataset tickets.jsonl --threshold 0.95. Each case is scored field-by-field; a result below the threshold exits non-zero, blocking the deploy.

What's the best way to run LLM evals in CI?

The best way to run LLM evals in CI is to treat evaluation like a unit test: run the function against a fixed dataset, set a pass threshold, and let the exit code gate the deploy. In Functional AI, add fn eval classify --dataset tickets.jsonl --threshold 0.95 as a pipeline step — it exits non-zero below the threshold, so the CI job fails automatically and the bad version never ships.

How do I test an AI agent before releasing it to users?

Testing an AI agent before release means defining expected outcomes for a range of starting states, running the agent multiple times against each, and gating on both a quality score and a consistency figure. In Functional AI: write a JSONL dataset with input (starting state) and expected (required outcome), then run fn eval classify --dataset agent-cases.jsonl --runs 3. An agent that scores well but produces different results on each run is not production-ready — gate on both score and consistency before promoting.

Get access

Functional AI is in private beta. If you want fn eval gating your own pipeline, join the waitlist — or become a design partner and shape the eval reports before they ship.