Xiaomi's model jumps near the frontier with just 30 RL steps, no plateau in sight
teortaxesTex · x · 2026-09-22
Commenting on Xiaomi's training details, nrehiew and others are surprised the model improved dramatically with only 30 RL steps — far fewer than expected — and question why training stopped given no visible plateau.
More from Models
- Amassing a PhD team is exactly what OpenAI did, dev notes in AI research debate — felpix_ · 2026-09-22
- Dev argues prompting alone can't get AI to solve natural science problems — felpix_ · 2026-09-22
- Dev claims benchmarks are 'absolutely meaningless' — models only differ by vibe — gnukeith · 2026-09-22
- Report: OpenAI Trained New Math Model in ~11 Days, Solved Longstanding Open Problems — felpix_ · 2026-09-22
- Jev Hallucinates at Rates Similar to Other Models, Developer's Examples Show — JeremyNguyenPhD · 2026-09-22
- Aikido Security releases Altar-1, its first open-weight security model — HankYeomans · 2026-09-22