Looped MoE Tuning: 2x Experts, 0.5x Looped Layers, 2x Loops, Attention Untied
burny_tech · x · 2026-09-30
Researcher rosinality proposes a more efficient setup for the looped MoE architecture: 2x experts, 0.5x looped layers, 2x loops, and attention untying.
The key claim is that with this configuration, the looped transformer problem shifts from being about inductive biases to being about better allocation of resources — reframing the architecture question as a compute-allocation question.
More from Research
- Will neural nets decompose into evolved symbolic systems? Researchers debate — QuintinPope5 · 2026-09-30
- Braco compresses visual tokens 144x at 95.2% accuracy with ~36% end-to-end speedup — Rui Zhong · 2026-09-30
- CaptchaArena releases 50K verified CAPTCHA trajectories; 9B agent hits 71.7% Pass@1 — ColumbiaUniversity · 2026-09-30
- Google's TabFM-Auto: LLM agent evolves data pipelines, lifting TabFM by 228 Elo and topping MLE-Bench — google · 2026-09-30
- PrismQuant: null-space rotations make INT4 near-lossless, only 0.22pp below FP16 on Llama-70B — NanyangTechnologicalUniversity · 2026-09-30
- LEGO-Anything: coding agents write Blender code to rebuild editable 3D scenes, 62.7% gains — AWS · 2026-09-30