Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs

simonguozirui · x · 2026-09-22

peanoai announced full-parameter reinforcement learning on TPUs, including MiMo-V2.6 at 310B and other stable training runs of 1,000+ steps across 1,000+ TPUs.

Key details: built on JAX so scaling up is a config change rather than a rewrite; optimized vLLM inference for faster rollouts; full bitwise trainer–sampler agreement in validation; trainer and sampler share one TPU ICI fabric, transferring all 310B parameters in under 2 seconds.

The result demonstrates a viable engineering path for very large-scale RL training on the TPU stack.

Original post →

More from Infra

Infra channel →