Theorem says Lean-verified AI sandboxes are months away, at 1-30KB of proofs verified per hour
ctjlewis · x · 2026-09-23
Theorem co-founders Rajashree Agrawal and Jason Gross explain why fully verified AI sandboxes are still months away: models are finally good enough to prove sandbox properties, but verification throughput is the bottleneck.
- Gross benchmarks against the recent Navier-Stokes result: models wrote 600,000 lines of Lean in 17 hours, or roughly 1KB-30KB verified per hour.
- A minimal sandboxable Linux is about 5MB, so full verification remains a handful of months away at current rates.
- Token cost is a major limiting factor.
- Theorem aims to ship a fully verified sandbox by the end of the year to stop agent escapes.
More from coding & agent
- I tested 8 AI phone call agents: platform vs. managed are two very different products — Evening_Hawk_7470 · 2026-09-23
- Building a lead-scoring pipeline in n8n — deliberately without an AI agent — FlakyBeyond5850 · 2026-09-23
- Building an EU-sovereign agent, IONOS cloud locks account after signup bug — tobowers · 2026-09-23
- Developer finds Luna 6 is a 'massive downgrade' from Luna 5.6 despite better benchmarks — skilliard7 · 2026-09-23
- Graph Engineering: Building Reliable AI Agent Systems as Explicit Task Graphs — Pavan_Belagatti · 2026-09-23
- Have Your Coding Agent Attach Flame Graphs to Every PR It Opens — DanielLockyer · 2026-09-23