A five-step recipe: train gpt-oss-120b with MiMo RL envs and slime to match Luna

andrew_n_carr · x · 2026-09-30

andrewncarr outlines a five-step open-source route: take the new RL environments from MiMo, combine with gpt-oss-120b and the slime training framework, and train the model to match Luna's performance.

Original post →

More from coding & agent

coding & agent channel →