A five-step recipe: train gpt-oss-120b with MiMo RL envs and slime to match Luna
andrew_n_carr · x · 2026-09-30
andrewncarr outlines a five-step open-source route: take the new RL environments from MiMo, combine with gpt-oss-120b and the slime training framework, and train the model to match Luna's performance.
More from coding & agent
- Using AI plan mode to audit an SEO workflow and find where to plug in Jev — gaganghotra_ · 2026-09-30
- Dev's AI workflow grew so complex he asked the model to explain his own system — gaganghotra_ · 2026-09-30
- Herdr review: a TUI built for juggling multiple Codex and Claude Code sessions — leebase65 · 2026-09-30
- OpenAI exec shouts out Pidot at DevDay, signaling an open ecosystem play — omarsar0 · 2026-09-30
- AMD's Hyperloom and ROCm 10: AI agents tune GPU kernels overnight with accuracy checks — AnushElangovan · 2026-09-30
- A Prompting Trick: Make Coding Agents State Their Views on Good Software First — nptacek · 2026-09-30