Engineering
Deep dives on verified inference: the architecture, the cryptography, the replay, and the parts we got wrong first. Written for the engineers who evaluate us.
Subscribe via RSS-
Committed computation traces, explained for engineers
Re-running a sample of the work catches a cheat one time in five. Proving every operation cryptographically still costs a large multiple of the inference on production-scale models. There is a middle path: make the executor commit to its intermediate activations before it knows what will be checked, then verify a handful of inter-layer constraints instead of re-running end to end. For a 24-layer transformer with five sampled constraints that is 86.8 percent detection at the same sampling budget, and the honest boundary is that it only beats replay against an adversary who cheats to save work.
Read article → -
Sampling as deterrence: the economics of unpredictable re-execution
Re-running every job to verify it doubles the compute bill, and full cryptographic proofs still cost a large multiple of the inference itself. Sampling sits between them: re-run an unpredictable fraction and let the economics do the rest. Cheating stops paying as soon as the penalty on detection exceeds the saving divided by the sampling rate, and across a relationship of any length there is no cheating policy that survives. The whole argument rests on one property, which is that the executor cannot predict which jobs get checked.
Read article → -
Why bit-exact determinism fails on GPUs, and why a tolerance-based comparison is the right check
If you verify AI inference by re-running it on independent hardware, the two results will not match bit for bit, even when both parties are honest. Floating-point addition is not associative, and batch composition, kernel selection and GPU generation all change the reduction order. Demanding bit-exact agreement would mean surrendering the independence that made the second run worth anything. A calibrated tolerance separates honest drift from model substitution by orders of magnitude.
Read article → -
Three Properties of an AI Evidence Record
An AI evidence record can have three distinct properties, each established by a different mechanism and each failing its own way: integrity since checkpoint (what witnessed transparency logs deliver), completeness (a process property, established from outside), and truth at the time of writing (which the log cannot give you). A log faithfully pins whatever it was handed, so the source of truth is a separate question from the log that holds the record.
Read article → -
Hardware Attestation Versus Independent Re-execution: Two Threat Models
Attestation reports and independent re-execution both get sold as verification, but they assume different adversaries. Hardware attestation is strong when the machine is intact and the adversary is software; it excludes physical attacks, and TEE.fail showed that a party with physical access can forge a report for work that never ran in an enclave. Re-execution assumes nothing about the executor's hardware, but costs compute and offers no confidentiality. Which one you need depends on whether your adversary is the operator.
Read article → -
What a Receipt Proves, and What It Does Not
A walk through the artifact Cyberian issues for each inference: its fields, the three checks a third party can run offline (output digest, Merkle inclusion, attestor signature), what passing them establishes (integrity and attribution), what re-execution adds (truth, to the assurance level), and the exact boundary of what a receipt does not prove.
Read article →