For the 11th session of Proof is in the Pudding, we teamed up with Archetype to talk about zkML. Running an LLM is already expensive. How much harder is it to prove that you ran it?
We start with what a proof of model inference actually guarantees, how inference differs from training, and the alternatives to using ZK. Then we open up a transformer and look at the computation we would have to prove.
From there, we discuss sumcheck, GKR, lookup arguments, and provers specialized for particular models. If you'd like to work through one of those building blocks yourself, our sumcheck tutorial includes code and exercises. We also covered GKR from a security angle in Session 3.
The second half looks at quantization, models designed to be easier to prove, and opportunities around KV caching. We finish with scaling, where verifiable ML might be useful, and a more speculative possibility: agents agreeing on circuits and exchanging proofs as they interact.
Jump to a topic
- 1:44: What is zkML?
- 5:32: Inference vs. training
- 7:52: Alternatives to ZK for verifiable inference
- 10:44: Inside a transformer
- 15:25: Proving transformer computation
- 27:12: Quantization and SNARK-friendly models
- 31:58: KV caching and optimization opportunities
- 35:08: Agent-to-agent proofs and dynamic circuits
- 37:18: How zkML scales
- 38:50: The market for verifiable ML
You can catch up on our previous session on Groth16 or browse the full series. Have a topic you'd like us to cover next? Let us know on Twitter/X!