A 1T-parameter MoE model reportedly learned math reasoning with zero-RL and no human solutions

i_dg23 · x · 2026-07-27

A new paper claims a 1T-parameter MoE model learned math reasoning with zero-RL, without showing any human-written solution traces.

Setup

Reported results

The image also highlights infrastructure optimizations such as mixed-precision control and context-parallel optimization, plus emergent behaviors like anthropomorphism, structured format, parallel reasoning, and context anxiety.

Original post →

More from Research

Research channel →