Products Demo Docs Blog Engineering About Contact Sign in Sign up
Engineering · · Philippe Laporte

Committed computation traces, explained for engineers

Re-running a sample of the work catches a cheat one time in five. Proving every operation cryptographically still costs a large multiple of the inference on production-scale models. There is a middle path: make the executor commit to its intermediate computation before it knows what will be checked, then check a few internal consistency constraints instead of re-running end to end. Here is how it works, what it buys, and where it stops.

This is a plain-language walk through a preprint I released on Zenodo in September 2026, "Committed Computation Traces: Hybrid Verification of Neural Network Inference Between Probabilistic Replay and Zero-Knowledge Proofs." The paper has the formal statements, the adversary taxonomy and the cost tables. This piece has the intuition. One caveat up front: what follows is research. The verification records we issue in production today come from sampled independent re-execution, and the construction described here is a direction, not a shipped feature.

The gap the paper is about

There are two established ways to check that an outsourced neural network inference was done correctly.

Probabilistic replay re-runs a random sample of the work on independent hardware and compares. With a fifth of the work re-run, a dishonest job is caught with probability one in five, at a fifth of the inference cost in overhead. The assurance is black-box: the verifier learns whether the output matches, and nothing about where a discrepancy came from.

Full cryptographic proof systems prove every arithmetic constraint of the forward pass with cryptographic soundness. A verifier can be nearly certain the computation was done exactly as claimed. The price, for models with hundreds of millions of parameters, has been a large multiple of the inference itself, often cited in the thousands when the preprint was written. Newer claims suggest that is falling fast; until they are independently reproduced, the gap below remains the practical one.

Between one-in-five at low cost and near-certainty at a high cost there is a lot of empty space. The question the paper asks is whether a verifier can get substantially more than one-in-five detection without paying substantially more than replay.

The idea: commit to the trace, then check constraints

A layered neural network is a chain. The input goes into layer one, layer one's output goes into layer two, and so on to the output. Write h_l for the activation vector coming out of layer l, and g_l for the function layer l computes, residual connections included. Then the whole forward pass is a set of L constraints:

h_l  =  g_l( h_(l-1) )        for every l = 1 .. L

A correct execution satisfies all of them. Now suppose the executor, instead of just handing over the final output, also commits to every intermediate activation before anyone tells it what will be checked. The commitment is a hash-based structure over per-layer digests, so it is small, cheap to produce, and binding: once committed, the executor cannot change a single activation without the change being detectable.

After the commitment is in, the verifier issues a challenge that selects k of the L constraints. For each selected layer l, the executor opens the two activations h_(l-1) and h_l against its commitment, and the verifier recomputes g_l on h_(l-1) itself and checks that it gets h_l. That is one layer's worth of computation per selected constraint, not a full forward pass, and the verifier never has to re-run the whole model.

The challenge has to be unpredictable to the executor and fixed after the commitment. In the paper it is derived from a secret key and the commitment itself, so the executor cannot know which layers will be opened until it has already bound itself to all of them.

Why checking five layers out of twenty-four catches most cheats

Here is the part that makes the construction worth more than its sampling rate suggests.

Suppose the executor falsifies the computation at layer l. Maybe it skipped the real work at that layer and wrote something plausible into the trace. What happens to the constraints downstream? Layer l's constraint is violated, because h_l is not what g_l would have produced. But layer l plus one's constraint is also violated, because the real g_(l+1) applied to the falsified h_l does not give whatever the executor wrote for h_(l+1), unless the executor went to the trouble of honestly computing layer l plus one on its own falsified input. And so on to the end. A single falsification cascades into a run of violated constraints from layer l to layer L.

The verifier does not need to sample the exact layer where the cheat happened. It needs to sample any layer in the violated run. The paper calls this the cascading constraint property, and it is what separates this construction from plain replay.

The arithmetic follows directly. If the cheat is at layer l and the verifier samples k layers uniformly out of L, the cheat escapes only if all k samples land in the first l minus 1 layers. Averaging over where the cheat might be:

escape, cheat at layer l   =   C(l - 1, k) / C(L, k)

averaged over l            =   (L - k) / ((k + 1) * L)

detection                  =   1 - (L - k) / ((k + 1) * L)

L = 24, k = 5              =   1 - 19/144   =   0.868      (86.8 percent)

Plain replay at the same sampling budget, checking five out of twenty-four units of work, detects a cheat 20 percent of the time. Same budget, 4.34 times the detection.

The honest boundary

Now the part I want to state as carefully as the paper does, because it is the difference between a result and a sales pitch.

The cascade only happens if the executor is lazy downstream of its cheat. An adversary who falsifies h_l and then honestly recomputes every later layer from the falsified value leaves exactly one violated constraint, at layer l. Against that adversary the verifier catches the cheat only if layer l is among its k samples, which is k / L, or about 20.8 percent for the same parameters. That is comparable to replay, not better than it.

So the construction does not raise detection against every adversary. It raises detection against the adversary who cheats to save work, which is the adversary who has an economic reason to cheat in the first place. An adversary who recomputes every downstream layer honestly has done most of the work the honest path would have required; what they have bought with their single falsification is a wrong answer they paid nearly full price for. The paper gives an explicit taxonomy of adversaries and states the detection probability for each, rather than quoting the best case.

Two variants

The construction above is the executor-side variant. It requires the executor to run instrumented code that captures and commits activations, which is a reasonable ask when the executor is a cooperating provider and an unreasonable one when it is a black box.

The paper also gives a prover-side variant for the black-box case. There, an independent prover re-executes sampled chunks of the work, as in ordinary replay, and while doing so captures and commits to its own intermediate activations. The detection probability is unchanged from replay, because the executor's internals are still invisible, but the verification record now carries an auditable trace of the reference computation path. That has forensic value: months later, a third party can see not just that the reference output differed from the executor's, but what the reference computation looked like layer by layer. It does not change how often a cheat is caught. The paper says so.

Costs, concretely

The commitments are hash-based Merkle structures over per-layer activation digests. Producing them adds hashing on the executor's side, which is small next to the matrix multiplications of a forward pass. The verifier's work per sampled constraint is one layer of computation plus checking two openings against the commitment. Hash-based commitments are also resistant to quantum attacks, which matters for records that are supposed to remain valid for years.

The paper analyses the numbers for a concrete model, BAAI/bge-large-en-v1.5, a 335-million-parameter, 24-layer encoder widely used for embeddings, and works through commitment sizes, opening sizes and verifier compute for that case. Polynomial commitment schemes are discussed as an optional upgrade path toward smaller proofs and algebraic verification, not as the base design. The base design is deliberately the cheap one.

How it relates to other work

Committing to intermediate activations is not a new idea on its own. TOPLOC, from Prime Intellect, commits to the top activations of intermediate layers through locality-sensitive hashing so that a verifier can check an inference provider's claims, and it is the closest neighbour to this construction. Freivalds' algorithm, from 1977, is the classic demonstration that a matrix product can be checked probabilistically far more cheaply than it can be computed. Sampled re-execution on independent hardware appears in several recent systems, Proof of Sampling among them. The paper positions committed traces as a bridge: the detection behaviour of constraint checking with the deployment simplicity of replay, and a clear account of which adversaries it does and does not improve against.

What I would want an engineer to take away

Three things.

Commit before challenge is the whole trick. Everything the construction buys comes from binding the executor to its full trace before it can learn what will be examined.

Cascading violations are why a small sample goes far, and only against the adversary who is trying to save work. State that boundary every time, because a construction that only reports its best case will be found out by the first careful reader.

And the cheap version is enough. Hash-based commitments and one layer of recomputation per sampled constraint get you the detection numbers above. The fancier machinery is an upgrade, not a prerequisite.


Disclosure: I am the author of the preprint, and I build inference verification at Cyberian Systems. The preprint is available on Zenodo (DOI 10.5281/zenodo.22880798). The production verification we ship today is sampled independent re-execution; the construction described here is research and is on our roadmap, not in the product.

PL
Philippe Laporte
Founder and CEO of Cyberian Systems, building verified AI inference infrastructure for regulated industries.

Read the docs · Try the live demo · RSS