// blog
Writing
Notes on building, evaluating, and shipping Agentic Functions.
- 5 min read
What Is Agent Orchestration?
Agent orchestration is the control layer that decides which AI agents run, in what order, with what context, and how the system recovers when any step fails.
- 7 min read
How Much Does an AI Agent Cost to Run in Production?
Concrete production cost breakdown for AI agents — real model prices, token math, tool call fees, and three levers that cut your inference bill in production.
- 6 min read
The Future of LLM Hallucination in Production
Hallucination rates don't improve linearly with model size. Here's where the problem is heading as agents compound errors — and how to gate against it.
- 6 min read
RAG Tools: A Production Engineer's Stack Checklist
The engineering guide to RAG tools: chunking strategies, hybrid retrieval, faithfulness thresholds, and CI gates that catch regressions before they ship.
- 6 min read
Eval Metrics: What to Measure, What to Gate, What to Skip
Which eval metrics to run in CI, how to calibrate thresholds, and the agentic failure modes most guides skip.
- 7 min read
How to Test Prompts: Assertions, Thresholds, and CI Gates
Write prompt tests that catch real regressions: test case structure, scorer selection, golden datasets, and five failure modes every team hits.
- 6 min read
Agent Loop: Architecture and Failure Modes
An agent loop is the perceive-plan-act cycle at the core of every AI agent. Architecture, two dominant patterns, and four production failure modes.
- 6 min read
AI Observability: A Production Engineer's Guide
AI observability captures traces, eval scores, and cost signals from agents at runtime—catching failures that HTTP 200 hides. A concrete engineering guide.
- 6 min read
Non-Deterministic LLMs: Own the Variance in Production
LLMs are non-deterministic even at temperature=0. Learn why batch-invariance failure is the real cause, and how to declare and gate on a variance budget in CI.
- 7 min read
LLM as Judge: How It Works and When to Use It
LLM as judge uses one model to score another's outputs. Learn how it works, the three judging modes, common biases, and how to wire it into your CI/CD pipeline.
- 6 min read
What is an agent runtime?
An agent runtime is the hosted layer that runs AI agents in production — executing the loop, gating changes with evals, and handling model fallback.
- 6 min read
What is harness engineering?
Harness engineering is the discipline of building the scaffolding that turns a model into a dependable agent — tools, context, evals, and guardrails.
- 8 min read
What Is an Eval?
An eval is an automated test that measures whether your AI prompt or agent does what you intend—dataset, task, scorer—and how to gate CI/CD on the result.
- 6 min read
AI regression testing in CI/CD
Four changes reliably break AI features: prompt edits, model swaps, provider drift, and schema changes. Here's the CI test suite that catches all four.
- 7 min read
What is an LLM evaluation platform?
An LLM evaluation platform runs automated quality checks on AI outputs—offline in CI and online in production—to catch regressions before they reach users.
- 6 min read
How to evaluate an LLM prompt before shipping it
Evaluating an LLM prompt means running it against a fixed dataset and blocking the deploy when the pass rate falls below your threshold. Here's how.
- 6 min read
How to add automatic model fallback to your AI app
Automatic model fallback switches your AI feature to a backup the moment the primary fails. Here’s the failure modes, the DIY path, and the runtime approach.
- 6 min read
The best way to host an AI feature as a function
The best way to host an AI feature as a function is an AI function runtime: versioning, eval gates, model fallback, and tracing without rebuilding the shell.
- 6 min read
LLM cost optimization in production
LLM cost optimization means running on the cheapest model that clears your quality bar. Here’s how quality-floor model selection works in production.
- 5 min read
What is prompt versioning?
Prompt versioning tracks every change to an LLM prompt—instructions, model, parameters, schema—as an immutable snapshot you can evaluate, compare, and roll back.
- 7 min read
The Eight Pillars of a Production-Grade Agent
The gap between a demo agent and a production agent is eight operational pillars wide. What they are, and how a runtime provides the shell so you don't rebuild it per Agentic Function.
- 6 min read
Functional AI: Agentic Functions, Not Prompts
Why we built a runtime that treats agents as versioned Agentic Functions instead of prompts copied between branches.