SOMA caps gradient variance for zero-order optimization, hitting SOTA on pretraining
StanfordAILab · x · 2026-10-01
François Chaubard's team unveiled SOMA (Sharded Optimization Mixture of Assemblies), a new architecture designed for zero-order (ZO) optimization that achieves SOTA for ZO methods on pretraining at controlled compute and parameter counts.
- The problem: ZO methods fail to improve loss as models scale because relative gradient variance grows linearly with the number of perturbed parameters.
- Key insight: Prior work innovates on the optimizer while keeping backprop-designed architectures; SOMA instead designs the architecture with ZO in mind, capping gradient variance as model size grows.
- How it works: The model is sharded into tiny experts, each trained independently on disaggregated GPUs with no communication of gradients, activations, or optimizer state during training.
- Inspiration: Loosely modeled on the thalamus and cortical columns, with each expert specializing on a subset of the training data.
More from Research
- Shibaura Tech's Ozaki Scheme II lands in CUDA 13.4, squeezing FP64 from AI-focused GPUs — udmrzn · 2026-10-01
- Functional Gradient Descent Beats Neural Nets by ~10x, New Paper Fixes Its Convergence Flaw — burkov · 2026-10-01
- "Mode-Hopping" in LLM Pretraining: OLMo3-32B Swings 81%→0%→81.7% in 40B Tokens — jiaxinwen22 · 2026-10-01
- Interleaved Head Attention Accepted at NeurIPS, Boosts RULER Retrieval 10-20% — _arohan_ · 2026-10-01
- BrainWorks 2026 workshop at MICCAI to tackle AI and data scaling laws in brain disease mapping — PTenigma · 2026-10-01
- LibraryDesignBench tests whether AI agents can design and effectively use code libraries — a1zhang · 2026-10-01