How do you verify an untrusted GPU host actually ran the model? Gonka's design notes on 3 cheating vectors

autoimago · reddit · 2026-09-17

A contributor to open-source project Gonka published design notes on verifying that untrusted GPU hosts in a decentralized inference network actually ran the paid-for open-weight models (DeepSeek V4-Flash, GLM-5.3-Flash, MiniMax M2.7), when re-running every request would double network costs.

Three cheating vectors and defenses:

Problems posed to the community: what output tolerance can't be gamed across GPUs/kernels; whether checks should be cost-weighted given MoE routing and 400K-token prompts; and whether sampling suffices or some requests need TEEs/ZK proofs. Code is public (inference-chain/), with HackerOne bounties for security breaks.

Original post →

More from Infra

Infra channel →