Two vLLM bugs hid in plain sight: Mamba state cache, decode-before-prefill and a 32-bit wrap

AI Engineer · youtube · 2026-09-20

AI21 engineers Asaf Gardin and Yuval Belfer walk through two vLLM bugs, both rooted in the Mamba state cache.

Case 1: gibberish once in 1,000 prompts

Case 2: logprob spikes every 12th step, mistaken for RL instability

Takeaway: stateful inference doesn't fail loudly — it lies with confidence. Both bugs surfaced under memory pressure and were found via logprob comparison.

Original post →

More from coding & agent

coding & agent channel →