Motif 3 Technical Report: A 314B Parameter Sparse MoE Model
Motif-Technologies · hf · 2026-08-11
Motif Technologies releases the technical report for Motif 3, a decoder-only Mixture-of-Experts (MoE) language model with 314B total parameters and 13.2B activated per token.
- Architecture: Each sparse MoE layer contains 384 routed experts (8 selected per token). It introduces Grouped Differential Latent Attention (GDLA), integrating grouped differential attention with compressed KV representations. Also features modified manifold-constrained hyper-connections and multi-token prediction.
- Training: Pretrained on approx. 12.5 trillion tokens across web, STEM, code, math, and multilingual content. Selective MXFP8 computation and memory-efficient fused kernels enable training with context lengths up to 256K.
- Post-training & Performance: The pipeline combines SFT, six RL-trained specialist teachers, and Multi-teacher On-Policy Distillation. The unified model demonstrates competitive performance against leading open-weight models, particularly in long-horizon agentic tasks, math reasoning, and long-context understanding.
More from Models
- Local Open-Source Model Self-Checks and Runs Code: What Do Subscriptions Still Buy? — truecakesnake · 2026-08-11
- OpenAI Accused of Stifling Model Creativity via Strict Temperature Controls — RileyRalmuto · 2026-08-11
- GLM-5.2 Pricing Drops: High-Level Intelligence Gets Dramatically Cheaper — tobowers · 2026-08-11
- Developer Praises Kimi K3: 'Like if Llama 4 Behemoth Was Real' — willccbb · 2026-08-11
- Alibaba Confirms Qwen3.8-27B Open Weights Landing This Week — max_paperclips · 2026-08-11
- Abacus AI Releases Smaug-Agentic, Topping Open-Source Leaderboard for Agentic Coding — bindureddy · 2026-08-11