// blog
Non-Deterministic LLMs: Own the Variance in Production
LLMs are non-deterministic even at temperature=0. Learn why batch-invariance failure is the real cause, and how to declare and gate on a variance budget in CI.
A non-deterministic LLM produces different outputs from the same prompt on successive calls — even at temperature=0, even with a fixed seed. This is not a defect: it is a measurable property of every hosted inference endpoint, driven by batch-composition changes in GPU kernels that shift floating-point reduction order. You cannot eliminate this variance on a hosted API. You can measure it, declare an explicit budget for it, and gate your deploys on it. That is the engineering response.
Why temperature=0 doesn't fix it
The popular fix — set temperature=0 — removes sampling randomness but leaves the deeper source untouched. Temperature controls how the decoder picks between tokens; at T=0 it picks the argmax. That part is deterministic on paper.
The logits feeding that argmax are not.
Research from Thinking Machines Lab identified the true mechanism: hosted inference servers process requests in dynamically sized batches, and the GPU kernels running RMSNorm, matrix multiplication, and attention are not batch-invariant. When batch size changes — depending on how many other requests are in flight at that moment — the internal floating-point reduction order changes. Floating-point addition is not associative: (a + b) + c ≠ a + (b + c). Different reduction orders produce different logits. A single flipped token in autoregressive decoding changes every subsequent token.
The Thinking Machines team ran 1,000 identical completions of the same prompt at temperature=0 on a production endpoint. Without batch-invariant kernels, they received between 80 and 100 distinct completions — variance caused by server load, not by sampling.
A 2025 study published at ACL tested six LLMs across eight benchmark tasks with maximally deterministic settings and found accuracy variations of up to 10% across five identical runs. No prompt changed. No model changed.
OpenAI exposes a seed parameter and a system_fingerprint field for tracking backend state. The API documentation is explicit: "Determinism is not guaranteed." As of mid-2026, both fields are marked deprecated in OpenAI's published OpenAPI spec. They were always best-effort controls.
How non-determinism breaks CI — and nobody notices
The trap plays out like this:
- You ship a new prompt version for a ticket classifier.
- CI runs the eval once. 97 of 100 samples pass. ✅ Green build. Merge.
- In production, variance is real. On unlucky batch compositions, the pass rate drops to 88%.
- Nobody notices until users complain.
The problem is not the model. It is that you tested a non-deterministic system as though it were deterministic. A single eval run is a single draw from a distribution. It proves nothing about the distribution.
The fix is N rollouts with a pass-rate gate:
fn eval classify@v3 \
--dataset tickets.jsonl \
--rollouts 10 \
--pass-rate 0.95 \
--threshold 0.8
This runs each sample ten times and requires 95% of rollouts to meet the 0.8 quality threshold. Declare this contract once in fn.yml and every deploy to Functional AI enforces it automatically:
function: classify
evals:
- dataset: tickets.jsonl
rollouts: 10
pass_rate: 0.95
threshold: 0.8
The contract travels with the function. If the model provider's serving infrastructure shifts — even silently — and your variance expands past the declared budget, the deploy fails before it reaches users.
The model fallback trap
Automatic model fallback keeps your agent running when a primary model is rate-limited or unavailable. But it introduces a second-order non-determinism problem: the fallback model has a different variance profile than your primary.
Your eval was green against gpt-4o-2024-08-06. In production, 3% of calls hit the fallback and run on a different model with a different output distribution. The function's live behavior no longer matches what you tested.
The fix is to eval your Agentic Function against every model in its fallback chain before shipping:
fn eval classify@v3 \
--dataset tickets.jsonl \
--rollouts 10 \
--pass-rate 0.95 \
--eval-fallbacks
Functional AI runs the eval against the primary and each configured fallback in sequence. If any model in the chain fails the variance gate, the deploy is blocked. Trust your agents in production means trusting the whole call path — including the path you hope never runs.
Measuring your variance budget
Before you declare a budget, measure one. Run a calibration eval at high rollout count:
fn eval classify@v3 \
--dataset tickets.jsonl \
--rollouts 50 \
--report variance
classify@v3 variance report
samples: 100
rollouts/sample: 50
mean pass rate: 0.964
p10 pass rate: 0.921
p5 pass rate: 0.884
worst rollout: 0.810
recommended gate: --pass-rate 0.90 --rollouts 10
Set your CI gate at p10 minus a small margin. The mean pass rate always looks fine — p10 is your honest signal. A two-point margin (p10 - 0.02) absorbs normal week-to-week drift in model provider infrastructure without triggering false positives.
Who owns it?
The classic ownership failure: eval scores drop 8 points between Monday and Tuesday. Prompt engineers blame the model. Infra blames a vLLM version bump. Nobody has the data to prove either side, so nobody fixes it.
An Agentic Function has a declared eval contract that runs on every version bump — prompt, model, or infrastructure dependency. If a new model snapshot changes variance enough to breach the gate, the deploy fails with a logged rollout distribution. The failing version is recorded. The responsible change is identified.
This is what engineering ownership looks like for non-deterministic systems: not eliminating variance, but making the variance budget explicit, versioned, and enforced on every deploy.
Self-hosted: the extra lever
If you run your own inference stack, one option unavailable on hosted APIs is batch-invariant kernels. vLLM ships this as VLLM_BATCH_INVARIANT=1, making inference results independent of batch composition. It requires NVIDIA GPUs at compute capability 8.0 or higher (A100, H100) and carries roughly a 20% throughput cost. For compliance use cases requiring bitwise reproducibility, that tradeoff is worth taking. For most production agents, measuring and gating on a variance budget is the more practical path.
FAQ
What does non-deterministic mean for an LLM?
A non-deterministic LLM returns different outputs for the same input across successive calls — same prompt, same parameters, different tokens. The cause: hosted inference servers run requests in variable-sized batches, and GPU kernels are not batch-invariant. Batch-size changes shift floating-point reduction order in matrix multiplications and normalization layers, producing different logits even at temperature=0.
Does temperature=0 make LLM output deterministic?
No. Temperature=0 makes the decoder pick the highest-probability token at each step, but the logits feeding that decision vary with batch composition on any hosted API. Research shows that 1,000 identical prompts at temperature=0 can produce 80–100 distinct completions due to batch-composition changes in the serving infrastructure.
How should I test a non-deterministic system in CI?
Run each eval case 5–10 times (rollouts) and gate on a pass rate, not a single pass/fail. Measure your p10 pass rate at high rollout count (50 rollouts), then set your CI gate at roughly p10 - 0.02. This catches real regressions while absorbing normal infrastructure drift.
What happens to variance when a model fallback triggers?
The fallback model has a different output distribution than the primary. If you only eval against the primary, you have no contract on what the fallback produces. Eval every model in the fallback chain before shipping.
Can I achieve bitwise determinism in production?
Only by self-hosting with batch-invariant kernels enabled (e.g., VLLM_BATCH_INVARIANT=1), which requires specific GPU hardware and carries roughly a 20% throughput cost. For most production agents, declaring and enforcing a variance budget is the practical alternative.
Ready to give your Agentic Functions a declared variance contract? Join the waitlist at fnai.dev/#waitlist.