Functional AI: Agentic Functions, Not Prompts
June 25, 2026 · 6 min read · Functional AI
Most teams ship their first agent the same way: a prompt in a string, a model call behind it, and a hope that nobody touches it. It works in the demo. Then it goes to production, and the cracks show.
Someone tweaks the wording to fix one edge case and silently regresses three others. There is no record of what the prompt looked like last Tuesday, no way to compare two versions, and no number anyone trusts that says whether today's change is better or worse. The prompt is the most important code in the codebase, and it is the least engineered.
A prompt is not a unit of software
Functions have an interface. They have inputs and outputs with types. You can test them, version them, diff them, and roll them back. None of that is true of a prompt sitting in a string literal.
So we stopped treating prompts as the unit and made the Agentic Function
the unit instead. In Functional AI, you write an fn.yml — a declarative file
that names the inputs, the outputs, the model, and how the thing should be
evaluated:
name: classify
version: 3
model: claude-opus-4-8
output:
category: string
priority: enum[low, high]
confidence: number
eval:
dataset: ./tickets.jsonl
threshold: 0.95
Compile that and you get classify@v3: an immutable, content-addressed Agentic
Function with a typed API. The prompt is an implementation detail the runtime
owns, not a string you babysit.
Versions are the whole point
Because every change produces a new immutable version, the questions that used to be unanswerable become trivial:
- What changed between v2 and v3? —
fn diff classify@v2 classify@v3 - Is v3 actually better? — the eval score, compared automatically to v2
- Production is on fire, undo it —
fn rollback classify --to v2, instantly
Rollback is instant because nothing is rebuilt. v2 still exists; you are only moving a pointer.
Evaluation is a gate, not a vibe
The reason prompt changes are scary is that nobody knows if they worked. We made evaluation a first-class command that exits non-zero below a threshold, so it drops straight into CI:
fn eval classify --dataset tickets.jsonl --threshold 0.95
If the score regresses, the build fails and the deploy never happens. A 2% drop that would have slipped past a human reviewer fails the pipeline instead.
Not another agent framework
We are deliberate about the words we use. Functional AI is not another agent framework you assemble. Your agent becomes an Agentic Function: inputs in, structured output out, the same way every time, with the metadata you need to bill and trace it.
That constraint is the feature. An Agentic Function you can test is one you can trust in production — and trust is the thing that has been missing from agents all along.
Ready to build one? Start with the Quickstart.
// private beta
Ship agents you can trust.
Build your prompt, agent, or multi-agent workflow as an Agentic Function. Functional AI hosts it, certifies each version before it ships, switches to a backup model when the primary fails, and lets you release without a redeploy. You just call it from your code.
Shipping agents today? Become a design partner — design partners shape the roadmap and get early access.