Products Demo Docs Blog About Contact Sign in Sign up
Engineering · · Philippe Laporte

Merkle commitments for batch inference receipts

A batch of a million embeddings should not need a million signatures, and an auditor who asks about one item should not need the other 999,999. A data structure from 1979 solves both problems, and it is the reason a verification record can be small, checkable offline, and durable. Here is what it does and, just as important, what it does not.

When an AI system processes work in batches, the evidence problem changes shape. Nobody wants a separate signed document for every one of a million vectors. Nobody wants to ship a million-item file to a regulator who asked about item 4,127. And nobody should have to trust that a file they were handed later is the same one that existed at the time. Merkle trees, introduced by Ralph Merkle in his 1979 doctoral thesis and used everywhere from version control to certificate transparency, handle all three.

Hashes first

A cryptographic hash function takes any input and produces a fixed-size digest, typically 32 bytes. Two properties matter here. It is one-way: you cannot recover the input from the digest. And it is collision-resistant: nobody can find two different inputs with the same digest. The practical consequence is that a digest is a binding fingerprint. If I give you the digest of a file today, and you obtain a file tomorrow whose digest matches, you know it is the same file to the byte.

That already solves one problem. A receipt can carry digests of the input, the output and the model weights rather than the data itself, and anyone holding the data can check that it matches. What it does not solve is scale and selectivity: a batch of a million items would need a million digests in the receipt, and checking one item would still require the whole list.

The tree

Take the million items. Compute a digest for each: in an inference batch, a natural leaf is the digest of the input, the digest of the output, an identifier of the model, and the item's position, hashed together. Now pair the leaves up and hash each pair to get half a million digests. Pair those and hash again. Keep going until one digest remains. That single 32-byte value is the Merkle root, and it commits to every leaf beneath it.

batch of 1,000,000 items

leaf     =  hash( model id, input digest, output digest, position )
parent   =  hash( left child, right child )
root     =  one value, 32 bytes, whatever the batch size

inclusion proof for item 4,127
         =  20 sibling digests  =  640 bytes
   check:   hash the leaf, fold in each sibling in order,
            compare the result with the root in the signed receipt

Change one byte of one leaf and its digest changes, which changes the digest of its parent, and so on up to the root. The root is therefore a fingerprint of the entire batch, and it is the same size whether the batch has ten items or ten million.

A receipt for the batch carries the root, the model identifier, the batch size, and a signature over all of it from whoever is vouching for the batch. One signature, one root, any number of items.

Inclusion proofs, or why the auditor does not need the whole batch

Here is the part that makes the structure useful rather than merely compact.

To prove that item 4,127 is in the batch committed by a given root, you do not need the other items. You need the item's own leaf, plus the sibling digest at each level of the tree on the path from that leaf to the root. For a million items that path is twenty levels deep, so the proof is twenty digests, about 640 bytes. The auditor hashes the leaf, combines it with the first sibling, hashes again, combines with the next, and after twenty steps arrives at a root. If it matches the root in the signed receipt, the item was in the batch. If it does not, either the item or the proof has been altered.

Notice what the auditor needed: the item itself, 640 bytes of proof, the signed receipt, and the signer's public key. No network access, no database, no cooperation from the party that produced the batch. That is what "checkable offline" means in practice, and it is why a record built this way can still be checked years later by someone who was not a party to the original transaction.

What the root does and does not prove

This is the section I most want people to read, because Merkle trees get credited with things they do not do.

The root proves membership and integrity. It proves that a given item was part of the committed batch, and that neither the item nor the batch has changed since the root was signed. If a party later produces a different output for item 4,127 and claims it was the original, the inclusion proof will fail against the signed root.

The root does not prove correctness. It binds what was claimed; it does not check the claim. A batch of wrong answers has a perfectly good Merkle root. Whether the outputs in the batch are actually what the named model produces on the named inputs is a separate question, answered by a separate mechanism: independent re-execution of a sample, or a proof of the computation, or a commitment to the intermediate trace. The tree is the container for the evidence, not the evidence itself.

The root also does not prove completeness on its own. It commits to what the producer chose to include. A producer who never records a job has nothing in the tree to be caught by. Completeness comes from elsewhere: from the consumer holding its own list of what was submitted and checking that each item has an inclusion proof, or from an append-only log that a third party witnesses.

Keep these three apart, membership and integrity, correctness, and completeness, and most confusion about verification records disappears.

Why hash-based, and why it ages well

A Merkle commitment needs only a hash function. No elliptic curves, no pairings, no trusted setup. That has two consequences worth stating.

It is cheap. Hashing a million leaves and building the tree is a fraction of a second of ordinary CPU time, negligible next to the inference that produced the batch.

It is durable. Hash functions of the kind used here are not known to be weakened by quantum computers in the way that most public-key schemes are, which is why hash-based signatures are the conservative choice in post-quantum standards: NIST's FIPS 205 standardises a stateless hash-based signature scheme for exactly that reason. A record whose integrity rests on a Merkle root over a modern hash is one you can expect to still be checkable when the original hardware, vendor and contract are all gone.

The signature over the root is a different matter. Whatever signature scheme is used to sign the receipt is what has to stay strong, and a durable design either uses a post-quantum signature or is built so the signature can be replaced without touching the tree.

Design details that matter in practice

A few things the textbook version leaves out.

Domain separation. Leaves and internal nodes should be hashed with distinct prefixes, so that an attacker cannot present an internal node as if it were a leaf. This is a well-known pitfall, and it is a one-line fix.

Position in the leaf. Including the item's index in the leaf pins each item to its slot and prevents two identical inputs from becoming ambiguous.

Odd sizes. When a level has an odd number of nodes, the tree needs a fixed rule for the leftover, and the same rule has to be used by everyone verifying. Certificate transparency's log format, specified in RFC 6962, is a good reference for a rule that has been argued over thoroughly, and it pins the leaf and node prefixes too.

A worked mental model

Think of the signed receipt as a sealed envelope with a fingerprint on the outside. The fingerprint is the root. The batch is the contents. An inclusion proof is a way to show a stranger that a particular page was inside the envelope when it was sealed, without opening the envelope or showing them any other page. And whether the page says something true is a question the envelope cannot answer, only the process that wrote the page can.

Once that model is in place, reading any verification vendor's claims gets easier. Ask what is in the leaf, who signs the root, what proves the leaf's contents are correct, and what proves the batch is complete. Four questions, four different mechanisms, and a tree that only ever answers the first.


Disclosure: I build inference verification at Cyberian Systems, and our batch receipts use Merkle commitments over per-item digests. The correctness of those receipts comes from independent re-execution of a sample, which is a separate mechanism from the tree and is covered in its own pieces.

PL
Philippe Laporte
Founder and CEO of Cyberian Systems, building verified AI inference infrastructure for regulated industries.

Read the docs · Try the live demo · RSS