TUM's Continuous Depth Batching Unlocks 99% of Speedup for Looped LMs
TUM · hf · 2026-09-28
A TUM team introduces the first efficient batching method for depth-adaptive inference in looped language models.
- Problem: looped LMs reuse shared layers a variable number of times (less compute for easy tokens), but tokens with different loop counts can't share a forward pass, breaking standard batching like vLLM
- Method: Continuous Depth Batching (CDB) forms new batches between loop steps, dynamically schedules looped/non-looped parts, manages looped KV-caching, and predicts token exits in advance to prepare batches asynchronously
- Results: experiments on Ouro 1.4B and Huginn 3.5B show fully looped architectures suit depth-adaptive inference best, while large non-shared layers (embeddings, LM head) complicate scheduling. CDB achieves up to 99% of the estimated maximum speedup, with further gains limited mainly by architecture and exit behavior
More from Research
- Cyber Index Alliance Launches With IBM, NVIDIA and Vercel to Standardize AI Cyber Defense Eval — ArtificialAnlys · 2026-09-28
- Artificial Analysis Details Cyber Index Methodology: Safety Refusals Score Zero, Tracked Separately — ArtificialAnlys · 2026-09-28
- Nissenbaum paper takes on privacy nihilism as AI inference erodes data-category frameworks — AllThingsApx · 2026-09-28
- Princeton researchers warn AI could slow science despite exploding paper output — AllThingsApx · 2026-09-28
- One-arm VR intervention during bimanual DAgger feels like an AI brain chip — neurosp1ke · 2026-09-28
- NeurIPS-rejected paper shows six agent dimensions to measure after benchmark saturation — random_walker · 2026-09-28