Products Demo Docs Blog Engineering About Contact Sign in Sign up
Engineering · · Philippe Laporte

Why bit-exact determinism fails on GPUs, and why a tolerance-based comparison is the right check

If you verify AI inference by re-running it somewhere else, the two results will not match bit for bit, even when both parties are honest. That is not a bug in the method. It is arithmetic. Here is what actually varies, and how to build a check that still catches a cheat.

The naive design for verifying outsourced inference is appealing: run the model again on a machine you control, and compare the bytes. If they match, the executor was honest. If they differ, it was not.

It does not work, and the reason is worth understanding precisely, because the failure mode tells you what the replacement check has to do.

Floating-point addition is not associative

Everything starts here. In IEEE floating point, (a + b) + c and a + (b + c) are not guaranteed to be equal. Each addition rounds to the nearest representable value, and where the rounding happens depends on the order of operations. The classic demonstration is (0.1 + 1e20) − 1e20, which gives 0, against 0.1 + (1e20 − 1e20), which gives 0.1. Sum a thousand random float32 values in chunks of 256 and then in chunks of 250, and the totals differ in the last digits.

A neural network forward pass is dominated by large reductions: dot products in matrix multiplications, sums in normalisation layers, softmax denominators in attention. A GPU computes each of these by splitting the work across thousands of threads and combining partial results. The combination order is a property of the kernel that ran, not of the mathematics. Two different kernels, both correct, will in general disagree in the low bits.

Where the variation really comes from

The folk explanation is that GPU threads finish in unpredictable order, so every run differs. That is mostly wrong, and the correction matters for anyone designing a verifier. Thinking Machines Lab published a careful analysis in September 2025 showing that a typical forward pass has almost no truly nondeterministic operations in it: for a fixed input shape, fixed batch composition and fixed kernel, running the same pass twice on the same machine yields bit-identical output. The instability comes from elsewhere.

It comes from batch invariance, or rather its absence. Many GPU kernels choose their reduction strategy based on the shape of the whole batch. Process one request alone and the kernel splits the work one way; process it alongside thirty-one others and the kernel splits it another way. The individual result changes in the low bits even though nothing about that request changed. On a production server, batch composition depends on who else is sending traffic at that moment, so from any one client's point of view the numerics are effectively random. The fix that analysis proposed, batch-invariant kernels, buys back run-to-run determinism on a single controlled stack at a real cost in throughput.

It also comes from everything a verifier deliberately does not share with the executor. Different GPU generations use different tensor-core datapaths. Different library versions select different kernels for the same operation. Autotuning picks a kernel based on timing measurements taken on that machine. Fused attention kernels tile the problem differently. TF32 and mixed-precision settings change how many bits survive each multiply. Any one of these changes the reduction order, and any change in reduction order moves the low bits.

Why the verifier cannot just match the executor's stack

You could, in principle, pin everything: the same GPU model, the same driver, the same library versions, the same batch shape, batch-invariant kernels throughout. Inside a single organisation that is worth doing for reproducibility. For verification it defeats the purpose.

The value of an independent verifier comes from its independence. It runs on hardware the executor does not control, in an administrative domain the executor cannot reach, and ideally on a different provider altogether so that no single vendor's failure compromises both sides. That independence is precisely what guarantees the stacks will differ. Demanding bit-exact agreement across independent stacks is demanding that both parties surrender the property that made the second run worth anything.

So the design question is not how to make two GPUs agree to the bit. It is how to build a comparison that tolerates honest numerical drift while still refusing everything an executor might do to cheat.

What honest drift looks like, and what cheating looks like

Take embedding models, which is where I have the most measured experience. Run the same text through the same model on two different GPU stacks and compare the output vectors by cosine similarity. Honest disagreement lives far out in the decimal places: the two vectors point in the same direction to within a tiny angle. Now run a different model, a smaller or older one that an executor might substitute to save money. The vectors point in visibly different directions. The gap between honest drift and substitution is not a matter of a few decimals; it is orders of magnitude.

That gap is what makes a tolerance-based check sound: a threshold that sits inside it rejects every wrong model and accepts every honest run, with margin on both sides. How the threshold is chosen and maintained is an implementation matter I will not go into here. The point of this piece is that the gap exists, and that it is wide.

The cheats a tolerance check catches

Model substitution. A cheaper model produces outputs that miss the threshold by a wide margin. This is the cheat with the biggest payoff for a dishonest executor, and it is the easiest to catch.

Skipped work. Random or cached outputs miss by even more. An executor that returns yesterday's result for today's input fails the moment the verifier re-runs today's input.

Silent hardware faults. A flipped bit in a weight tensor or a failing memory module produces outputs that drift far outside honest variation. This is not fraud, but it is exactly what a buyer wants to know about.

Quantisation downgrades deserve honesty. Running the same architecture at lower precision moves the outputs by less than substitution does, sometimes by not much more than honest drift. Whether a tolerance check catches it depends on how aggressive the quantisation is and how tight the calibrated threshold is. This is the one case where the margin is narrower, and where anyone using this method should measure rather than assume.

What a tolerance check cannot do

It cannot prove the executor ran the exact bytes of the model, because that is not a well-defined claim across two GPUs. What it proves is that the output is consistent with the reference model to within a tolerance that no cheaper computation has been observed to satisfy.

It cannot, on its own, catch an adversary who runs the real model and then perturbs the output by an amount smaller than the tolerance. The reason this is not a serious hole is economic: to land within tolerance, the adversary has to run the real model on the real input, at which point they have paid the full cost of honesty and gained nothing by tampering except a slightly worse answer they could have delivered for free.

And it says nothing about jobs the verifier did not re-run. Sampling is a separate design question, with its own economics, and I treat it in a separate piece.

Why this is the right check, not a compromise

Engineers coming from cryptography sometimes hear "tolerance" and think "weakened." I would put it the other way. A bit-exact check across independent hardware is not a stronger guarantee; it is an unattainable one, and any system that claims it is either running both sides on the same stack, which gives up independence, or has quietly defined "exact" as "within the noise of a single vendor's kernels."

A calibrated tolerance is the check that matches the physics of the hardware. It is measurable, it is auditable, because the calibration data can be shown, and it separates the honest case from every dishonest case worth an executor's while by a margin you can state. That is what a verification method is supposed to do.


Disclosure: I build verification by independent re-execution at Cyberian Systems, and the comparison method described here is the one our receipts rest on for embedding workloads. The analysis of batch invariance is Thinking Machines Lab's, from "Defeating Nondeterminism in LLM Inference," September 2025.

PL
Philippe Laporte
Founder and CEO of Cyberian Systems, building verified AI inference infrastructure for regulated industries.

Read the docs · Try the live demo · RSS