Developer proposes mixture-of-attention architecture with centroid clustering for token routing
InfamousTrouble7993 · reddit · 2026-08-18
A developer released LogPose, a mixture-of-attention LM architecture addressing attention dilution and dead experts. It uses a differentiable soft K-means clustering router to route tokens to specialized attention heads, plus GRU-based recurrent routing and linear routing. The implementation includes a Llama-style decoder baseline (RMSNorm, RoPE, SwiGLU, GQA), built-in GSM8K and MBPP benchmarks, KV and routing-state caching, and MLflow integration. The author asks if this is novel, fearing prior work.
More from Research
- AURORA-LM: Diffusion Models Master High-Fidelity Text Representations — jiqizhixin · 2026-08-18
- Training model without data inspection? The model trains you — kalomaze · 2026-08-18
- Study Finds 3.8M Agent Skill Files Across 282K GitHub Repos — dair_ai · 2026-08-18
- 300-Page Monograph: Engineering Reliable Coding Agents as Systems — heyneighbor · 2026-08-18
- Tsinghua and ByteDance's CUDA Agent Writes Better CUDA Than Human Experts — anselm · 2026-08-18
- Paper on "Machine Studying" Explores New Paradigm for AI Agents — lateinteraction · 2026-08-18