Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs
simonguozirui · x · 2026-09-22
peanoai announced full-parameter reinforcement learning on TPUs, including MiMo-V2.6 at 310B and other stable training runs of 1,000+ steps across 1,000+ TPUs.
Key details: built on JAX so scaling up is a config change rather than a rewrite; optimized vLLM inference for faster rollouts; full bitwise trainer–sampler agreement in validation; trainer and sampler share one TPU ICI fabric, transferring all 310B parameters in under 2 seconds.
The result demonstrates a viable engineering path for very large-scale RL training on the TPU stack.
More from Infra
- Sentdex benchmarks openjev: 169ms on Dell GB10 vs 137ms on RTX 3090 — Sentdex · 2026-09-22
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- Underdog's Husky Inference Engine Claims 4.5x Speedup Over MLX, 730 tok/s on MacBook — jimmykoppel · 2026-09-22
- Cloudflare Python Workers go generally available after two-year preview — Simon Willison · 2026-09-22
- Fighting AI crawler traffic: beyond Turnstile, Cloudflare's AI Labyrinth as an option — fforres · 2026-09-22
- Raspberry Pi locks devices to original RAM size, blocking aftermarket memory upgrades — ngxson · 2026-09-22