Temperature Zero Is Not Deterministic, and Your Debugging Assumes It Is

Share
Temperature Zero Is Not Deterministic, and Your Debugging Assumes It Is. Abstract deep dive illustration in orange and dark grey on debugly.dev

The bug report said the feature "sometimes" returned a different classification for the same input, and the engineer's first move was to pin the temperature to zero, because the folklore says zero means deterministic. It did not. Two identical requests, minutes apart, same model, same parameters, produced different outputs, and the reproduction that relied on determinism collapsed. The assumption that temperature zero gives you a repro is one of the most expensive false beliefs in LLM engineering, and it breaks the most basic tool we have: the repro.

This is the mechanics of why identical prompts at temperature zero still differ, and what a debug process that cannot assume determinism has to look like instead.

This was a classification feature on a hosted model behind a load balanced inference fleet, observed across repeated identical requests, and the findings apply to any serving stack you do not fully control.

What temperature zero actually controls

Temperature scales the logits before sampling. At zero, sampling collapses to argmax, the single most probable token. That removes the randomness the setting governs, and only that randomness. It does not make the probability computation itself identical across runs, and it is the computation, not the sampler, that varies.

Argmax is only deterministic if the logits are bit identical. They are not, for several independent reasons, each of which is normal operation, not a fault.

The sources of variation with the sampler removed

Floating point non associativity across parallel reduction. The logits are produced by large matrix multiplications, summed in parallel. Floating point addition is not associative, so the order of summation changes the low order bits of the result, and the order depends on how the work is split, which depends on batch composition, sequence lengths and hardware threads. Two runs with the same prompt but different batch neighbours can produce logits that differ in the last bits.

Batching and dynamic padding. Inference servers batch requests to stay efficient. Your identical prompt lands in different batches with different other requests, different padding, different kernel tiling, and the arithmetic shifts in the low bits accordingly.

Kernel and hardware heterogeneity. A fleet has mixed GPU generations and driver versions, and the load balancer does not pin your request to one chip. Different hardware runs different kernels, and different kernels round differently.

Ties and near ties at argmax. When two candidate tokens have nearly equal probability, a low order bit flip changes which one argmax picks, and a one token difference early in the generation changes every subsequent context, so the outputs diverge visibly. The divergence is not noise in the sampler. It is chaos in the autoregressive loop, amplified from a bit.

Model and weight updates. The fleet may be mid rollout of a new checkpoint, and "the same model" is two models for the duration of the deploy, which is the prompt deploy problem in treat a prompt change like a deploy, one level down.

Why this wrecks the repro

The classical repro is: run the same input, observe the same output, bisect the change. Every step assumes the system under test is a function. An LLM behind a serving fleet is a distribution, and at temperature zero it is a distribution with a very narrow spread, not a point. So a repro that needs the model to repeat itself is a repro that fails intermittently, and the engineer who trusts the first run builds the fix on a sample of one.

This is the same lesson as flaky tests are telling you something, except the flake is not in your test, it is in the physics of your dependency.

What a nondeterminism native debug process looks like

Record everything that defines the request, not just the prompt. The model identifier and checkpoint, the exact parameters, the token ids if you can get them, the request id, and the timestamp. The request id lets the provider tell you which replica and checkpoint served it, which turns "it differed" into "it differed because it hit the other checkpoint", a fact instead of a mystery.

Treat outputs as samples, not values. For any behaviour you care about, run the input several times and record the spread before you call it a bug. A classification that flips one run in fifty at temperature zero is a boundary case in the model, and the fix is to move the input away from the boundary or to add a guard, not to rerun until it passes, which is the retry storm from reviewing retry and backoff logic applied to luck.

Pin what you can, and assume the rest varies. Some providers offer a seed and a pinned replica for debugging. Use them when available, and know that a seed pins the sampler, not the arithmetic, so it narrows the spread rather than removing it.

Make the product tolerant of the spread. The durable fix for boundary inputs is not determinism but design: confidence thresholds with a fallback path, structured output with validation and retry against a schema, and idempotent downstream effects so a differing answer does not become a differing side effect. Tolerance is the property you can actually guarantee.

The eval implication

Evals that run each case once are measuring a sample of the distribution and calling it the model. The same case run ten times gives you a variance, and the variance is information: cases with high variance at temperature zero are the boundary cases, and they are exactly the cases your product will fail on in production. The eval harness in testing LLM features that are non deterministic must therefore report variance per case, because a mean accuracy that hides a high variance case is the average latency fallacy in miniature.

The rule

Temperature zero removes the sampler's randomness and nothing else. The logits still vary with batching, hardware, kernels and checkpoints, and argmax over near ties turns bit level variation into visible divergence. Stop building repros on the belief that the model is a function. Record the full request identity, sample the spread, pin what the provider lets you, and design the product to tolerate the spread you cannot remove.

The repro is still your most powerful tool. It just has to be a statistical repro now, and the moment you accept that, the flake stops being noise and starts being the diagnosis.