Three properties of an AI evidence record
There is a growing category of products that promise a record of what an AI system did. Registries, audit trails, evidence packs, receipts. They use overlapping words, and a buyer reading three vendor pages in a row could be forgiven for thinking they compete on the same thing.
They mostly do not. An evidence record can have three distinct properties, each established by a different mechanism, and each with its own way of failing. Naming them separately is the single most useful thing a buyer can do before signing anything, and it is also the fairest way to read what any one vendor, including me, actually claims.
Property one: integrity since checkpoint
The first property is that a record has not been altered since some known moment. This is the property transparency logs deliver, and they deliver it well.
The pattern is a decade old. Certificate transparency, the system browsers use to catch mis-issued certificates, works by publishing every certificate into an append-only log built as a Merkle tree, with the log's head periodically signed and cross-checked by independent parties. Once an entry is in and the head is witnessed, removing or rewriting it would change the head, and anyone holding an earlier head would see the inconsistency. Sigstore's Rekor does the same for software signatures. The math is hash functions and signatures, the verification runs offline with public keys, and the trust rests on the witnesses being independent of the log operator.
The AI evidence world now has its own instance of this. TRACE is an open specification developed under the Linux Foundation, with profiles that draw on the IETF's attestation and transparency work, and its public registry is an append-only, checkpointed log whose checkpoints are countersigned by an independent witness. The registry's maintainers published a detailed account of the design, and it is worth reading for its candour as much as its content. It states plainly that the registry is built to be distrusted: anyone can verify the chain of checkpoints with public keys and open tooling, without taking the operator's word for anything.
That is exactly the right claim, and it is exactly one property. A witnessed log establishes that a record, once checkpointed, has not been changed. It says nothing about whether the record was complete, and nothing about whether it was true when it was written. The deep-dive concedes the first of those directly. The second is where most of the confusion in this category lives.
Property two: completeness
The second property is that every event that should have a record has one. The registry deep-dive puts the limit in one sentence: a party that never records something has nothing to witness. No log, however well witnessed, can prove the absence of an omission, because an omission leaves nothing in the log to examine.
Completeness is a process property, not a cryptographic one, and it is established from the outside. Three approaches work in practice.
The consumer keeps its own list. If the party relying on the records maintains an independent count of what it submitted, it can demand an inclusion proof for each item and notice the gaps. This is the cheapest control and the most often skipped.
Records carry sequence. If each record references the previous one, or carries a monotonically increasing counter that the consumer can see, a missing record leaves a visible hole. A hole is not proof of wrongdoing, but it is proof of incompleteness, which is what the property needs.
Both sides record. When an exchange has two parties, each keeping a record that references the other's, an omission on one side is contradicted by the other. This is the shape the IETF drafts on bilateral attestation are exploring, and it is the most robust option where it can be arranged.
None of these makes the log complete. They make incompleteness detectable, which is the honest version of the property.
Property three: truth at the time of writing
The third property is the one buyers are usually asking for when they say "prove what the AI did": that the record describes what actually happened. The model named in the record ran, on the input named, and produced the output named, with nothing substituted, skipped or adjusted.
A witnessed log cannot establish this. It faithfully pins whatever it was handed. If the record was false when it entered the log, the log guarantees that it will stay exactly as false forever, with an excellent timestamp. So the third property has to come from somewhere else, and there are four candidate sources, each with its own threat model.
The operator's own log. The party that ran the model writes down what it ran. This is the default, and it is an assertion. It is worth exactly the operator's word, which is fine until the operator's word is what is in dispute.
Hardware attestation. Confidential-computing hardware measures the code and weights loaded into a protected environment and signs a report that a remote party can check against the silicon vendor's certificate chain. This is strong evidence when the machine is physically intact and the adversary is software. It rests on two things nobody can verify remotely, the vendor's root of trust and the physical integrity of the machine, and the vendors themselves place physical attacks outside their threat model. Where a record's truth rests on attestation, an operator with physical access to the hardware can produce a false record that attests as genuine. I have written separately on that trade-off, and it is not an argument against attestation; it is an argument about which adversary it defends against.
Independent re-execution. A verifier on separate hardware, under separate administrative control, re-runs the work and signs a record binding the model, the input and the output. Its authority comes from the fact that it never ran the original job and never trusted the machine that did. It offers no confidentiality, since the verifier sees what it re-runs, and it tolerates the numerical drift between GPUs rather than demanding bit-exact agreement. This is the source I build, and it is the only one of the four whose guarantee survives an operator with physical control of the executor.
Cryptographic proofs of the computation. Prove every arithmetic step of the forward pass so a verifier can check the claim without re-running or trusting anything. This is the strongest source and, for models of any real size, the most expensive by three to four orders of magnitude today. It is where the field is heading, and it is not where most production evidence comes from yet. Constructions that sit between re-execution and full proofs, binding the executor to its intermediate computation and then checking a sample of its internal constraints, are the active middle ground.
The important thing is not which source is best. It is that the source of truth is a separate question from the log that holds the record, and a vendor who answers the first with the second has changed the subject.
How the three compose
Once the properties are separated, the architecture almost writes itself, and it is not a single product.
The log gives you integrity since checkpoint. Whatever goes in stays as it was, witnessed by parties who do not work for the operator. TRACE's registry is a good instance of this; so is any log built on the same pattern.
The record's contents give you truth, or fail to. A record produced by the operator carries an assertion. A record carrying a hardware attestation carries evidence that holds while the hardware holds. A record produced by independent re-execution carries evidence that holds whether the hardware held or not. The log does not care which; the buyer should.
The consumer gives you completeness, by keeping its own list and demanding a proof for every item on it.
Put together, a record produced by independent re-execution, anchored into a witnessed log, with the consumer holding its own submission list, delivers all three properties, and each property is checkable by a party who trusted none of the others. That is where I expect this category to settle, and anchoring re-executed records into open, witnessed logs is where I intend to take our own. The registry deep-dive explicitly invites any receipt producer to seal records and have them witnessed, which suggests its authors see the same composition.
Four questions for any vendor
Which of the three properties does your evidence establish, and which does it not?
What does the truth of a record rest on: the operator's word, the hardware, or a party who never ran the job?
How would I detect a record that was never written?
Can I verify a record offline, with public keys, without you?
A vendor with one property and the honesty to say so is worth more than a vendor claiming three. The registry deep-dive is a model of the first kind. Read it, then ask your other vendors the same questions it answers.
Disclosure: I build independent re-execution at Cyberian Systems, which is the third property in this piece. The comment I left on the registry deep-dive made the same point in shorter form, with my affiliation stated. Anchoring records into witnessed logs is a direction, not something we ship today.