112 bugs across 84 projects: LLMs pass proof-of-concept but fail developers' own tests

lulzxdxdxd · reddit · 2026-10-11

An arXiv paper (2610.10150) analyzes 112 bugs across 84 projects, finding LLM-generated code often passes proof-of-concept stages but fails tests written by the developers themselves.

Crucially, benchmark scores swing significantly based on evaluation design alone, suggesting current LLM coding benchmarks are poor predictors of whether code is actually correct — a methodological warning for anyone relying on SWE-bench-style leaderboards.

Original post →

More from coding & agent

coding & agent channel →