OpenAI's math gains likely from massive Lean-based RL environments; post-training is "rich man's inference"
yacineMTB · x · 2026-10-09
Ofir Press speculates OpenAI's math breakthroughs come from building many RL environments around Lean formalizations of math problems—starting with hand-crafted simpler ones, then letting models autonomously generate environments at scale from arXiv papers. Yacine reshared it with the quip that "post training is the rich man's inference."
More from Models
- Front-end aesthetic tasks still favor Claude and Kimi, blogger observes — vista8 · 2026-10-10
- "Living Weights" ships in TensorFold: local models update their own weights while in use — EAccelerate_42 · 2026-10-10
- Claude Code User Burns 25% of Weekly Quota Overnight; Says Opus 5.5 Leads a Generation in Creative Tasks — AlchainHust · 2026-10-10
- Google's Astra crushes Claude Opus 15-4 at chess despite being 360 Elo points weaker — pvncher · 2026-10-10
- GPT 6.1 Sol splits opinion: dubbed worst current model, yet some find it good value — dotey · 2026-10-10
- Multiple devs are already building dedicated inference engines for GLM-5.3-Flash — lxfater · 2026-10-10