AFM3 20B paper proposes prompt-level pruning that activates only 20% of MLP layers
Aaaaaaaaaeeeee · reddit · 2026-08-04
- The post points to an OpenReview paper on AFM3 20B and highlights its instruction-following pruning approach.
- The model is designed to activate only about 20% of active MLP layers, while also using MoE sparsity.
- A key idea is that it is trained from scratch to use the same experts per prompt rather than switching per token or per layer.
- The author argues this could make a 30B active MoE behave more like a 14B active model in memory-bandwidth terms, and notes a similarity to the recently discussed Session-Adaptive Orthogonal Distillation idea.
More from Research
- Revisiting Gwern's Scaling Hypothesis and the Core Value of Pretraining — BlancheMinerva · 2026-08-04
- Andriy Burkov shows a 100-page deep reinforcement learning book — burkov · 2026-08-04
- Berkeley paper turns Gemini Robotics On-Device into a humanoid specialist via CLIFT — berkeley_ai · 2026-08-04
- LLM evaluation research says small prompt changes can flip benchmark rankings — jindong_wang92 · 2026-08-04
- A verification skill every agent needs: computer and browser use — vikvang1 · 2026-08-04
- University of Michigan lab opens five AI, ECG and multi-omics research jobs — kevinnbass · 2026-08-04