What a DER Can Rerun: A Candidate Means of Compliance for ML in Aviation Certification
Every design assurance discipline in aviation rests on one property: evidence a second party can independently check, not evidence a first party attests to. DO-178C did not ask a supplier to promise its code was tested. It required structural coverage analysis, verification cases, and traceability that a certification authority or a Designated Engineering Representative could rerun and confirm for themselves.
What happens when the artifact is a trained model
A DER can rerun a structural coverage report. A DER cannot rerun a trained model the way they rerun a test case, and a supplier's statement that its ML component behaved correctly on a given input is not, by itself, something a certification authority can independently check. That gap is exactly what EUROCAE WG-114 and SAE G-34, working jointly, and EASA's MLEAP project are working through right now. The open question for anyone building toward a certification basis that includes an ML component: what counts as an acceptable means of compliance for the data-recording and continued-airworthiness objectives once the artifact under review is a model rather than a deterministic program.
A candidate answer, borrowed from a principle design assurance already follows
Design assurance has always separated the party that produces an artifact from the party that verifies it. Applied to ML inference, that separation means the system that ran a given input through a given model is architecturally distinct from the system that checks the result. A cryptographic receipt, generated per mathematically verified inference, lets a certification authority or a DER independently re-derive that a specific output came from a specific model on a specific input. That is a structural property of how the inference was executed and checked, not a procedural sign-off added on top of it.
What is actually live, named honestly
Today, that capability applies to embeddings workloads and to ONNX-exportable models, verified through a bring-your-own-ONNX upload. That is worth naming precisely because a meaningful share of the ML already entering avionics and autonomy certification bases, perception and sensor fusion in particular, ships as ONNX-exportable models. For that slice of what UAS autonomy, eVTOL, ATM, and defense-autonomy suppliers are building today, a cryptographic receipt per mathematically verified inference is something you can generate and hand to a DER now, not a roadmap item.
A separate conversation, kept separate
Some autonomy stacks are moving toward model architectures more complex than a single exportable graph. Where that is true, what evidence those components would need to produce for a certification basis is a design-partner conversation about a future candidate means of compliance. It is not a claim about what runs in production today, and we are not making one. It is a conversation worth having early, while WG-114 and MLEAP are still writing the rules that will define acceptable evidence, rather than after a certification basis is already fixed.
Who we are talking to
We are keeping a close watching brief on WG-114, SAE G-34, and MLEAP, and we want to talk with engineering and certification leads at UAS autonomy, eVTOL, ATM, and defense-autonomy suppliers, as well as with the DO-178C, DO-330, and ARP4754A consultancies that walk those suppliers through certification. If you sit on one of these working groups, or you are trying to work out what evidence a DER will actually accept for the ML component in your next certification basis, that is the conversation worth starting now.