TestPrism: single-reference test eval overstates quality by 2x, shows NJU-LINK

NJU-LINK · hf · 2026-10-09

NJU-LINK's TestPrism benchmark (300 tasks, 3,000 valid/invalid candidate implementations) shows coding-agent test generation scores only 28.00% on their Joint Success Function—requiring tests to fail the initial program, accept all valid and reject all invalid candidates—versus 59.67% under single-reference evaluation. Their TestHelix pipeline with peer cross-validation and recursive self-improvement lifts scores by 8.67–9.00 points.

Original post →

More from coding & agent

coding & agent channel →