Automated research agent finds silent numerical bug in FlashInfer kernels; fix PR accepted

timshi_ai · x · 2026-08-28

An automated research system scored a real win: Josh Tobin's team had previously built a reward-hacking judge for performance optimization tasks. While applying it to new autoresearch work, the system discovered that some FlashInfer kernels (the library underpinning vLLM and SGLang) used a hardcoded value of -50,000 as a masked-attention sentinel — even though valid QK values can be smaller.

Such silently-wrong corner cases are notoriously painful to find (recall the historical flash-attention debate). The system not only located the issue but submitted a fix PR, which human experts accepted.

Related event: Automated AI Research Agent Finds Numerical Bug in FlashInfer Kernels(3 posts)→

Original post →

More from coding & agent

coding & agent channel →