Hyperball Optimizer from Princeton & Stanford Breaks LLM Training Bottlenecks

burkov · x · 2026-08-14

A joint research team from Princeton, Stanford, and Tsinghua University has released Hyperball, an optimizer wrapper designed to address the diminishing performance gains of matrix-based optimizers like Muon as model scales increase.

Core Mechanism

Results

On Qwen3-style models up to 1.2 billion parameters, Muon combined with Hyperball demonstrated superior training performance.

Original post →

More from Research

Research channel →