TestPrism: frontier suites pass 59.67% of references, 28% of full panels

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

Han Li, Lingxiang Hu, Jiacheng Huang, Ziqian Jiang, Jingkai Luo, Wei Gao, Yunfan Tan, Zun Wang, Jiaheng Liu

cs.SE

2026-10-09

TestPrism grades each generated suite on 10 implementations. The best of six model families reaches 28.00% joint success, versus 59.67% when only the reference must pass.

What problem this solves

Generated tests are usually accepted on two checks: fail on the untouched program, pass on one reference solution. SWT-Bench and TDD-Bench Verified both score that way. A requirement often allows several correct implementations, and several patches that run while breaking the spec. A suite can lock onto how the reference is written and reject other valid code. It can also miss a bug the reference happens not to contain, and still be counted as a success.

TestPrism replaces that with one joint check. The panel is hidden while the suite is written. The suite must fail on the initial state, accept every valid implementation, and reject every invalid one.

Method

Seventeen public benchmarks are adapted into 2,000 tasks, split into a disjoint 1,700 for training and 300 for testing. The test split covers six languages and 67 repositories. Each task has ten implementations: the reference, four other valid ones, and five invalid ones, 1,500 valid and 1,500 invalid in all. Labels come from the source verifier. Every check passing means valid. Any failure means invalid. Labels do not move during evaluation.

Valid alternatives are not near-copies. Their mean AST Jaccard distance must be above 0.5, measured on depth-3 syntax subtrees in the edited region, with identifiers normalized so a renamed variable does not count as diversity. Invalid patches pass the diversity check when they fail different verifier checks. Matching failure sets must also differ in structure. From 15,148 source tasks, after irreproducible environments are removed, two experts independently review whether the verifier overfits the reference. Cohen's kappa before adjudication is 0.68. Three hundred tasks remain. The training split is fixed before these hidden panels are assembled.

Joint Success Function (JSF) averages 300 tasks. A task scores 1 only if the suite fails on the initial program, accepts every valid implementation, and rejects every invalid one. A missing run or an execution error scores 0. With only the reference left in the panel, JSF collapses to Reference Acceptance Rate (RAR), and it cannot exceed RAR. The gap is the share of tasks that pass the single-reference check and fail the joint one. Valid acceptance, invalid rejection, and their harmonic mean H are reported alongside. Change coverage, delta C, is the fraction of executable reference edits the suite hits. Entry coverage measures executable code in entry spans, and can extend past the diff.

TestHelix changes the writing loop. Isolated agents each produce a test paired with a repair. A pair is kept only when the tests fail reliably on the initial program and pass on that repair. Cross-execution then runs one agent's tests against another's repair, which is where overly strict assertions and incomplete repairs show up. A reviewer who never sees the repairs checks assertions against the public requirements, and every rejection needs public evidence. If several qualified repairs all pass, the loop searches for an input that separates them. The outer loop is recursive self-improvement (RSI): AutoSaddler, driven by Opus 4.6, edits the generation strategy on 600 training tasks. The search reward is RAR, because JSF is too sparse to credit small gains. At evaluation the strategy is frozen, and the pair count and per-task inference budget match the comparator.

Results

Fourteen configurations cover six model families. The table is the highest-JSF setting in each family. All 300 tasks have complete observations.

ModelRARValid acceptInvalid rejectJSF
Claude Fable 5.159.67%72.67%79.33%28.00%
Kimi K356.67%63.67%83.00%24.67%
Qwen3.8 Max54.67%77.00%66.00%23.67%
GPT 5.6 sol-max54.00%61.33%86.00%23.33%
DeepSeek V4 Pro41.00%57.67%70.33%13.33%
GLM 5.343.00%50.33%86.33%11.33%

Claude Fable 5.1 leads, and 28.00% joint success against 59.67% on the reference still leaves a 31.67 point gap of tasks the reference check accepts and the full panel does not. Five of the six families reject invalid patches more readily than they accept alternative valid ones. Qwen3.8 Max flips that pattern: 77.00% valid acceptance, 66.00% invalid rejection. GLM 5.3 has the highest rejection rate and the lowest JSF. Rejecting defects is not enough. Across all fourteen configurations, change coverage correlates with JSF at r = 0.91 and entry coverage at r = 0.96. Claude's change coverage is 82.33%; its entry coverage is 54.33%. Reaching the code does not mean the assertions judge it correctly.

Holding the suites fixed and growing the panel keeps stripping successes. On 50 tasks, with the six families weighted equally, mean JSF is 49.33% on the reference alone, 38.10% after five added valid implementations, 30.49% after five added invalid ones, and 24.33% after nine alternating additions. In this diagnostic, count and structural diversity move together.

The failure sample is 50 joint-failure trajectories from each family's best configuration. After 14 timeouts are dropped, 286 remain. Two experts labeled them, kappa 0.81. Faulty checks are 50.0%, coverage gaps 34.6%, and errors in the tests' own execution 15.4%. Missing inputs are 30.8% and wrong criteria 26.9%. A poker suite never compared an ace-low straight with three of a kind, so a patch that misread the straight still passed. A bracket parser checked only the output count, so a patch that dropped valid groups still passed. Both suites fail on the initial program and pass on the reference.

TestHelix is measured only on GPT 5.6 sol and Kimi K3. The per-task inference cap matches each model's native harness, Codex (max) and Kimi Code. Those native rows are the same two systems as in the table above.

SettingGPT 5.6 solKimi K3
Native agent23.33%24.67%
One pair, no cross-check20.33%21.67%
3 pairs, submit one at random21.00%22.67%
3 pairs + cross-validation27.67%29.00%
5 pairs + cross-validation28.33%29.67%
3 pairs + cross-validation + RSI32.33%33.33%

One pair sits 3.00 points under the native agent. Three pairs with a random submission add at most 1 point over one pair. Turning on cross-validation moves GPT from 21.00% to 27.67% and Kimi from 22.67% to 29.00%. A fifth pair adds 0.67. RSI adds another 4.33 to 4.67. Against the native harness the gain is 8.67 to 9.00 points. Against the one-pair run it is 11.67 to 12.00. The abstract uses the first baseline. The body uses the second. The one-pair loop is the weaker of the two. On Kimi, full TestHelix reaches 33.33% JSF while invalid rejection falls from 83.00% to 78.67% and H from 72.06% to 70.73%. Getting every label right on a task is a different quantity from the average rejection rate.

Why it matters

Shipping a suite because it fails on the initial program and passes on the reference will merge tests into CI that reject legal alternatives or miss bad patches. On the strongest configuration, more than half of the reference successes do not survive the panel.

With the inference budget matched, the points come from cross-checking repairs and from the strategy search. Moving from three pairs to five barely moves JSF. Kimi's 33.33% is higher than Claude's native 28.00%, and those runs were not on a shared budget. Joint success is still about one task in three.

Limitations

Valid and invalid are verifier outcomes. The appendix states that a patch counts as valid if every check passes and invalid if any check fails, regardless of what the patch was meant to do. A verifier overfit to the reference will penalize a suite that is more permissive and more correct. Kappa on the usability review is 0.68. The related-work discussion of the oracle problem makes the same point: seeing more behavior does not make the labels semantically complete.

In the panel-growth diagnostic, candidate count and AST diversity move together, so the drop cannot be pinned on either one alone. The error taxonomy covers 286 failed trajectories, not the suites that already passed. TestHelix was not run on Claude Fable 5.1, the strongest JSF baseline, and the strategy search optimizes RAR. Python accounts for 205 of 300 tasks. C++ and Go have 8 each. A nonzero exit can be a failed assertion or a broken test script, so a failure on the initial state is not always a meaningful one.

Terms

Source

Related papers

All paper explainers