Evaluating an LLM prompt before shipping
Evaluating an LLM prompt before shipping means running it against a labelled dataset and checking whether the output meets a defined quality threshold — so regressions are caught before users see them. Functional AI turns that check into a single CLI command: set a dataset and a threshold, and every version of your prompt, agent, or workflow is scored before it reaches production.
Evaluation in Functional AI is a command, not a notebook. You define a dataset and a threshold; Functional AI turns every change into a pass/fail signal.
Dataset format
A dataset is a JSONL file — one JSON object per line — with an input and an
expected field:
{"input": "ticket-001.pdf", "expected": {"category": "billing", "priority": "high"}}
{"input": "ticket-002.pdf", "expected": {"category": "support", "priority": "low"}}
expected may be a full output object or a partial one. Only the keys you
specify are scored, so you can assert on category alone while ignoring
confidence.
Scoring
fn eval classify --dataset tickets.jsonl --threshold 0.95
Each case is graded field-by-field. The reported score is the fraction of
asserted fields that matched. A score below --threshold exits non-zero.
Open-ended output
Field matching works when the output is structured. For free-text or agent output, exact match is the wrong tool: use model-graded checks instead. Treat those scores as a signal, not ground truth — LLM judges carry positional and self-preference bias and drift between runs. Functional AI tracks graded scores per version so drift is visible, but a human review cadence stays the calibration layer that keeps automated scoring honest.
Multi-run consistency
Models are non-deterministic. Use --runs to execute each case multiple times
and measure variance:
fn eval classify --dataset tickets.jsonl --runs 5
The report includes a consistency figure — how many of the N runs produced identical output — alongside the score.
Regression detection
Every eval is compared to the last recorded run for that Agentic Function. The report shows the delta versus the previous version, so a quiet 2% regression doesn't slip through.
Running LLM evals in CI
Because fn eval exits non-zero below threshold, gating a deploy is one step:
- name: Evaluate function
run: fn eval classify --dataset tickets.jsonl --threshold 0.95
If the eval fails, the job fails, and the deploy never runs. No custom failure logic, no post-deploy rollback — the bad version never reaches users.
Testing an AI agent before releasing it to users
An agent is harder to evaluate than a single prompt: it calls tools, spans multiple turns, and its output may vary across runs. The evaluation approach is the same — define what success looks like, run it against test cases, gate on a threshold — but consistency matters more.
Write test cases as a JSONL dataset where input describes the starting state
and expected captures the required outcome:
{"input": {"user_message": "cancel my subscription"}, "expected": {"action": "cancel", "confirmed": true}}
fn eval classify --dataset agent-cases.jsonl --threshold 0.90 --runs 3
Check both the score and the consistency figure. An agent that returns the correct answer two runs out of three is not production-ready. Gate the deploy when both metrics clear the threshold.
Functional AI tracks the score delta against the previous version, so a quiet regression between agent releases doesn't slip through undetected. To evaluate your own agents before they reach users, join the waitlist.
FAQ
How do I evaluate an LLM prompt before shipping it?
Evaluating an LLM prompt before shipping means running it against a labelled
dataset and checking whether the output meets a defined quality threshold — so
regressions are caught before users see them. In Functional AI: create a JSONL
file of input/expected pairs and run fn eval classify --dataset tickets.jsonl --threshold 0.95. Each case is scored field-by-field; a result below the
threshold exits non-zero, blocking the deploy.
What's the best way to run LLM evals in CI?
The best way to run LLM evals in CI is to treat evaluation like a unit test:
run the function against a fixed dataset, set a pass threshold, and let the exit
code gate the deploy. In Functional AI, add fn eval classify --dataset tickets.jsonl --threshold 0.95 as a pipeline step — it exits non-zero below
the threshold, so the CI job fails automatically and the bad version never ships.
How do I test an AI agent before releasing it to users?
Testing an AI agent before release means defining expected outcomes for a range
of starting states, running the agent multiple times against each, and gating on
both a quality score and a consistency figure. In Functional AI: write a JSONL
dataset with input (starting state) and expected (required outcome), then
run fn eval classify --dataset agent-cases.jsonl --runs 3. An agent that
scores well but produces different results on each run is not production-ready —
gate on both score and consistency before promoting.
Get access
Functional AI is in private beta. If you want fn eval gating your own
pipeline, join the waitlist — or
become a design partner and shape the eval reports before
they ship.