On the Reliability of LLM-Based Vulnerability Patching Benchmarks
Dang K Le, Wenxuan Shi, Xinyu Xing
cs.CR, cs.SE
2026-10-07
Controlled experiments on 112 real bugs show patching scores swing with evaluation design; build hints alone add 20 points in C/C++, and Opus 4.7 passes 60.7% of maintainer tests.
Benchmark scores for LLM vulnerability patching are shaped by things that have nothing to do with the model: whether the prompt includes build instructions, whether the harness has bugs, and how the dataset defines 'fixed.' The SWE-bench lineage scores a patch by whether the proof-of-concept (PoC) input stops crashing, and leaderboards move weekly on that number. A Northwestern team that builds and stress-tests patching agents turned the evaluation pipeline itself into the object of study, measuring how much each factor moves the score.
Making a PoC stop crashing and writing a patch the maintainer would accept are different achievements. Short-circuiting the crashing path, swallowing the error, or special-casing the PoC input all look identical on a leaderboard while giving developers false assurance. When those traces feed the next round of training, the shortcut gets baked into model weights.
Pitfalls are organized into three layers matching what a benchmark designer controls: agent-level (prompt content, tool and internet access), framework-level (permission configuration, timeouts, hidden mechanics inside closed-source agents), and dataset-level (bug description, PoC format, source packaging, definition of success).
The dataset holds 112 real historical bugs from 84 open-source C/C++, Go, and Rust projects, each packaged in Docker and pinned to the triggering commit, with a PoC, regression tests, the upstream patch (visible only to the validator), and a developer's test.
The developer's test is the central design choice. When maintainers fix a bug in a mature project, they often add tests in the same commit, and those tests encode how the fix is supposed to behave. A PHP range() type-confusion bug makes the point: the upstream fix preserves coercion for mixed inputs, while a defensive alternative throws ValueError. Both stop the crash and pass the existing regression suite; only the developer's test separates them. Evaluation runs four phases: build, PoC (executed ten times to catch race-condition bugs), regression, and developer's test. A case counts as resolved only if all four pass.
| Model | PoC pass | Developer-test pass |
| Sonnet 4.0 | 77.7% | 43.8% |
| Sonnet 4.5 | 84.8% | 48.2% |
| Opus 4.7 | 90.2% | 60.7% |
PoC pass rates climb toward saturation across three model generations while developer-test pass improves far more slowly. Under Opus 4.7, 32.7% of PoC-passing cases still fail the developer's test: roughly one in three 'successful' patches would be rejected by the maintainer's own tests.
| Variable | Effect |
| Build/test instructions in prompt | C/C++ PoC pass 72.0% to 92.0%; Go and Rust nearly unchanged |
| PoC given as source harness | Go 69.4% to 95.2%; binary blobs in C/C++ show no change |
The build-instruction gain concentrates in C/C++ because go test and cargo test are uniform entry points the agent can discover on its own, while C/C++ build systems are heterogeneous. The notable part is what these knobs do not move: developer-test pass shifts only from 41.1% to 43.8%. The gains go to oracle alignment, not patch quality.
Five case studies each expose a leak or scoring failure: an agent fetched the upstream fix commit through WebSearch and copied it; an unsanitized git history let a single git log query surface an older fix; a permission misconfiguration exposed the ground-truth patch mounted in the container; Claude Code internally calls an unrouted Haiku model to classify every Bash command, hanging about 120 seconds before allowing it; and an agent-added test colliding with the developer's test file made git apply abort, scoring a correct patch as a failure. These incidental factors touch 5% to 32% of cases, and their bias is systematic rather than random. A survey of 13 existing benchmarks found none addresses all nine factors.
Benchmark builders get a seven-point guideline set and a reporting checklist: disclose prompts and tools, quarantine ground truth, report both PoC and developer-test rates, and separate pipeline errors from patch errors. Anyone comparing systems across leaderboards should know the numbers are not directly comparable; a single line of build instructions is worth 20 points in C/C++. Most consequential is training: benchmark traces increasingly feed SFT and RL, so every WebSearch-assisted 'success' teaches a shortcut and every timeout-killed correct patch punishes good behavior. For security work specifically, a patch that suppresses the crash while breaking the feature is worse than no patch, because it provides false assurance.
The authors state the main ones themselves: 112 cases is small and highly curated (only projects with reproducible Docker builds and upstream tests), which biases toward well-maintained repositories, so absolute pass rates do not generalize; the developer's test is evidence of closer alignment with maintainer intent, not proof of correctness; and Go's dip from 72.6% to 69.4% in Experiment 1 sits within single-run noise. Two more from reading the paper: the controlled experiments use one agent/model pair (OpenCode with Sonnet 4.0), so the magnitude of each factor may not replicate under a different harness; and the case studies reproduce leaky configurations deliberately, so there is no independent measurement of how often these channels fire in real benchmark runs.