MiMo-V2.6 livestreams its RL run at ~2B tokens/step as Stanford's Marin pretrains in public
stanfordnlp · x · 2026-09-20
Stanford NLP spotlighted two open model experiments you can follow live:
- MiMo-V2.6's RL run: after 6 months of silence studying how far RL can scale, the team is scaling three things — compute (2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments/harnesses (multi-task agentic RL, mixed harnesses in one run), and grader compute (agentic in-group credit assignment with test-case and rubric-based rewards). Details will be open-sourced piece by piece; the run streams live.
- Marin's live pretraining: Percy Liang and the OpenAthena team are pretraining in public, letting outsiders follow along like team members.
More from Models
- System One Models Like Jev as Primitives for Frontier Agents — omarsar0 · 2026-09-20
- Code-only heuristic policies can beat frontier models on Craftax, evals researcher says — JoshPurtell · 2026-09-20
- Codex usage reset now live for all, big OpenAI release teased for Tuesday — kimmonismus · 2026-09-20
- Jev beats GPT-5.6 Luna on PR review: 1.93x faster at $0.0014 per run — aniketmaurya · 2026-09-20
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20