Meta researcher shares batch-ratio token dropping training trick
Meta researcher TimDarcet shared a training optimization: drop tokens by batch ratio instead of per-sample probability, keeping tensor shapes fixed and improving GPU efficiency, sparking discussion on applying similar pruning to full-layer computation.
2026-09-09 ~ 2026-09-09 · 3 related posts
- Meta researcher's training trick: drop a batch proportion instead of per-sample to keep GPU efficiency — TimDarcet · 2026-09-09
- Drop a fixed batch proportion instead of per-sample tokens: capi author shares training trick — giffmana · 2026-09-09
- Researchers weigh token-drop trick: variable-length sequences make it far harder — TimDarcet · 2026-09-09