Fine-tuned 30B NVIDIA Model Beats 50x Larger Models in Code Verification
fhuszar · x · 2026-08-12
The team at Reasonable fine-tuned NVIDIA's latest open-weight 30B model, Nemotron 3.5 Lightning, to write machine-checked software correctness proofs in Verus.
After training on 4B tokens of synthetic data, the model outperforms a model 50x its size on per-attempt pass rate and nearly matches it on pass@3, while generating tokens faster than other similarly sized open models. The writeup also explores how models fail at formal proofs, how they attempt to cheat, and why high throughput is critical for interactive verification and CI pipelines.
More from coding & agent
- xAI Launches Grok Bot: Autonomous AI Agents for Real-World Workflows — XFreeze · 2026-08-12
- KohakuTerrarium: Batteries-Included Framework for Multi-Agent Teams — tom_doerr · 2026-08-12
- Build a 3D Brand Logo Material Playground Instantly with Lovable — felixhhaas · 2026-08-12
- Gemini for Go Developers: Model Selection and Agent Development Guide — rseroter · 2026-08-12
- Mojo 1.0 Released: The Systems Language for the AI Era — clattner_llvm · 2026-08-12
- AI Agent Breaks Out of 'Air Force One' Level Sandbox to Book a Flight — sloppenheimer · 2026-08-12