Researchers weigh token-drop trick: variable-length sequences make it far harder
TimDarcet · x · 2026-09-09
Replying to a question about whether a tokendrop-style permute+slice trick could fully skip a layer, TimDarcet concedes it doesn't work with packed sequences: it would require variable-shape kernels and hit heterogeneous distributed compute — still interesting, but clearly harder than the fixed-seqlen case.
Related event: Meta researcher shares batch-ratio token dropping training trick(3 posts)→
More from Research
- Tabular foundation models over LLMs for predictions: talk at AI Engineer Paris — helloiamleonie · 2026-09-09
- Composing Continual Learning Mechanisms Boosts 100-Task Memorization Retention 28x to 34.9% — DanielKhashabi · 2026-09-09
- How OpenAI found a singularity in Navier-Stokes — and what it means for AI in science — eigensteve · 2026-09-09
- OpenAI's millennium proof dispute: fraud accusations, Altman denial, Tao's open science warning — The Decoder · 2026-09-09
- François Fleuret: we know training works, but not why — inductive bias, distillation and more remain opaque — francoisfleuret · 2026-09-09
- OpenWAM Releases Open Modular World-Action Model with Strong Real-Robot Performance — OpenWAM · 2026-09-09