112 bugs across 84 projects: LLMs pass proof-of-concept but fail developers' own tests
lulzxdxdxd · reddit · 2026-10-11
An arXiv paper (2610.10150) analyzes 112 bugs across 84 projects, finding LLM-generated code often passes proof-of-concept stages but fails tests written by the developers themselves.
Crucially, benchmark scores swing significantly based on evaluation design alone, suggesting current LLM coding benchmarks are poor predictors of whether code is actually correct — a methodological warning for anyone relying on SWE-bench-style leaderboards.
More from coding & agent
- Photocraft repo: clean-room Photoshop reimplementation in pure Rust hits 44.5k GitHub stars — victor_explore · 2026-10-11
- Microsoft Foundry demos automatic observability and optimization for enterprise agent fleets — AmyKateNicho · 2026-10-11
- Microsoft demos full agentic AI workflow inside VS Code with Foundry — AmyKateNicho · 2026-10-11
- Veracode: AI now writes half of all code, but security pass rate stuck at 56% — WeldPond · 2026-10-11
- Professor credits Claude with building a bespoke lualatex rig that unblocked a year-stalled book — prof_g · 2026-10-11
- One-month review: Devin SWE-2 + Opus + Codex multi-model coding workflow — CtrlAltDwayne · 2026-10-11