AI automated research finds numerical bug in vLLM/SGLang backend
josh_tobin_ · x · 2026-08-28
Josh Tobin shared a win for automated research using AI. Their reward hacking judge for performance optimization discovered an edge case in FlashInfer, the underlying library for vLLM and SGLang. The kernel used a hardcoded value of -50,000 as a masked-attention sentinel, even though valid QK values can be smaller. Such silent numerical errors can cause significant debugging headaches, demonstrating the potential of AI in finding deep code bugs.
More from Infra
- DFlash 2 Introduces Block-Diffusion Speculative Decoding to Speed Up GLM-5.3-Flash — songhan_mit · 2026-08-28
- Dev runs 125B Qwen model locally on M3 Max at 70 tok/s via MLX — mayfer · 2026-08-28
- NVIDIA launches Mesh open compute network to aggregate idle GPUs for AI — nvidia · 2026-08-28
- Hot Chips 2026: inference chips enter an "era of ferment" with divergent bets — BenBajarin · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28
- Analyst predicts CXL standard commercialization to break the memory wall — BenBajarin · 2026-08-28