Looped Model Architecture: 8B Params Outperform 32B in Reasoning
SonglinYang4 · x · 2026-08-01
This research delves into Looped Models, which reuse the same weights across depth to achieve a better compute-parameter trade-off.
The authors conducted rigorous apples-to-apples ablations (matching both training and inference FLOPs) across architectures from Ouro to Huginn. Huginn performs better overall, with major gains attributed to the "loop-in-the-middle" (sandwich) design and input injection.
Experiments show that an 8B-A0.8B Huginn MoE trained on 500B tokens approaches or surpasses a 32B-A3.2B feedforward MoE on reasoning benchmarks like GSM8K (83.6% vs. 80.8%), while using 75% fewer resident parameters under matched compute constraints.
More from Research
- Humanoid Robot Dodges 19/20 Thrown Balls Using Onboard Sensors — ChongZzZhang · 2026-08-01
- TMLR Adopts Fractional Authorship, Weighing Credit by 1/k per Author — thegautamkamath · 2026-08-01
- ACE-Data-0: A Large-Scale Multimodal Dataset for Embodied AI — liuziwei7 · 2026-08-01
- Waterloo's R2L Lab to Recruit PhDs, Focusing on Agents and Reasoning Research — hllo_wrld · 2026-08-01
- AgentIR: Deep Research Agents That Leverage Reasoning Context for Retrieval — hllo_wrld · 2026-08-01
- Research: Easy Model Identity Laundering via Simple Fine-Tuning — generativist · 2026-08-01