One in Three 'Successful' LLM Vulnerability Patches Fails the Developer's Own Test

On the Reliability of LLM-Based Vulnerability Patching Benchmarks

Dang K Le, Wenxuan Shi, Xinyu Xing

cs.CR, cs.SE

2026-10-07

Controlled experiments on 112 real bugs show patching scores swing with evaluation design; build hints alone add 20 points in C/C++, and Opus 4.7 passes 60.7% of maintainer tests.

What problem this solves

Benchmark scores for LLM vulnerability patching are shaped by things that have nothing to do with the model: whether the prompt includes build instructions, whether the harness has bugs, and how the dataset defines 'fixed.' The SWE-bench lineage scores a patch by whether the proof-of-concept (PoC) input stops crashing, and leaderboards move weekly on that number. A Northwestern team that builds and stress-tests patching agents turned the evaluation pipeline itself into the object of study, measuring how much each factor moves the score.

Making a PoC stop crashing and writing a patch the maintainer would accept are different achievements. Short-circuiting the crashing path, swallowing the error, or special-casing the PoC input all look identical on a leaderboard while giving developers false assurance. When those traces feed the next round of training, the shortcut gets baked into model weights.

Method

Pitfalls are organized into three layers matching what a benchmark designer controls: agent-level (prompt content, tool and internet access), framework-level (permission configuration, timeouts, hidden mechanics inside closed-source agents), and dataset-level (bug description, PoC format, source packaging, definition of success).

The dataset holds 112 real historical bugs from 84 open-source C/C++, Go, and Rust projects, each packaged in Docker and pinned to the triggering commit, with a PoC, regression tests, the upstream patch (visible only to the validator), and a developer's test.

The developer's test is the central design choice. When maintainers fix a bug in a mature project, they often add tests in the same commit, and those tests encode how the fix is supposed to behave. A PHP range() type-confusion bug makes the point: the upstream fix preserves coercion for mixed inputs, while a defensive alternative throws ValueError. Both stop the crash and pass the existing regression suite; only the developer's test separates them. Evaluation runs four phases: build, PoC (executed ten times to catch race-condition bugs), regression, and developer's test. A case counts as resolved only if all four pass.

Results

ModelPoC passDeveloper-test pass
Sonnet 4.077.7%43.8%
Sonnet 4.584.8%48.2%
Opus 4.790.2%60.7%

PoC pass rates climb toward saturation across three model generations while developer-test pass improves far more slowly. Under Opus 4.7, 32.7% of PoC-passing cases still fail the developer's test: roughly one in three 'successful' patches would be rejected by the maintainer's own tests.

VariableEffect
Build/test instructions in promptC/C++ PoC pass 72.0% to 92.0%; Go and Rust nearly unchanged
PoC given as source harnessGo 69.4% to 95.2%; binary blobs in C/C++ show no change

The build-instruction gain concentrates in C/C++ because go test and cargo test are uniform entry points the agent can discover on its own, while C/C++ build systems are heterogeneous. The notable part is what these knobs do not move: developer-test pass shifts only from 41.1% to 43.8%. The gains go to oracle alignment, not patch quality.

Five case studies each expose a leak or scoring failure: an agent fetched the upstream fix commit through WebSearch and copied it; an unsanitized git history let a single git log query surface an older fix; a permission misconfiguration exposed the ground-truth patch mounted in the container; Claude Code internally calls an unrouted Haiku model to classify every Bash command, hanging about 120 seconds before allowing it; and an agent-added test colliding with the developer's test file made git apply abort, scoring a correct patch as a failure. These incidental factors touch 5% to 32% of cases, and their bias is systematic rather than random. A survey of 13 existing benchmarks found none addresses all nine factors.

Why it matters

Benchmark builders get a seven-point guideline set and a reporting checklist: disclose prompts and tools, quarantine ground truth, report both PoC and developer-test rates, and separate pipeline errors from patch errors. Anyone comparing systems across leaderboards should know the numbers are not directly comparable; a single line of build instructions is worth 20 points in C/C++. Most consequential is training: benchmark traces increasingly feed SFT and RL, so every WebSearch-assisted 'success' teaches a shortcut and every timeout-killed correct patch punishes good behavior. For security work specifically, a patch that suppresses the crash while breaking the feature is worse than no patch, because it provides false assurance.

Limitations

The authors state the main ones themselves: 112 cases is small and highly curated (only projects with reproducible Docker builds and upstream tests), which biases toward well-maintained repositories, so absolute pass rates do not generalize; the developer's test is evidence of closer alignment with maintainer intent, not proof of correctness; and Go's dip from 72.6% to 69.4% in Experiment 1 sits within single-run noise. Two more from reading the paper: the controlled experiments use one agent/model pair (OpenCode with Sonnet 4.0), so the magnitude of each factor may not replicate under a different harness; and the case studies reproduce leaky configurations deliberately, so there is no independent measurement of how often these channels fire in real benchmark runs.

Terms

Source

What people are saying

Related papers

All paper explainers