Rewinding One Request in a Continuously Batched LLM Server to Block Banned Phrases
AlpinDale · x · 2026-09-08
AlpinDale's new post tackles a subtle serving problem: banning whole phrases (not tokens) from a model's streamed output. Banning token sequences directly has too large a blast radius, and by the time the engine detects a complete phrase, tokens may already be on the user's screen and baked into the KV cache.
- The solution: hold back output that could still complete a banned string; on a match, rewind that individual request and sample a different continuation.
- Implementing this inside an asynchronous scheduler is tricky because workers may already be executing later steps. The author built it in Sonar (formerly Aphrodite Engine).
- Prior art includes ExLlamaV3, KoboldCpp, and AntiSlop Sampler, but not in this async batched setting.
More from Infra
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11