Mixture-of-Depths routes FLOPs per token to cut transformer compute

burny_tech · x · 2026-09-23

A post highlights the paper "Mixture-of-Depths: Dynamically allocating compute in transformer-based language models" (arXiv:2404.02258).

Original post →

More from Research

Research channel →