TestPrism: single-reference test eval overstates quality by 2x, shows NJU-LINK
NJU-LINK · hf · 2026-10-09
NJU-LINK's TestPrism benchmark (300 tasks, 3,000 valid/invalid candidate implementations) shows coding-agent test generation scores only 28.00% on their Joint Success Function—requiring tests to fail the initial program, accept all valid and reject all invalid candidates—versus 59.67% under single-reference evaluation. Their TestHelix pipeline with peer cross-validation and recursive self-improvement lifts scores by 8.67–9.00 points.
More from coding & agent
- BrickBench benchmarks agentic LEGO design: agents pass constraints, trail humans — Peter Kulits · 2026-10-09
- autoicd-mcp ships automated ICD-10 medical coding MCP server with 74,000+ code search — modelcontextprotocol · 2026-10-09
- pohjola-api: agent-native Finnish company data API priced at $0.01 per call via x402 — modelcontextprotocol · 2026-10-09
- Dev recreates NES classic Faxanadu with Claude Opus, sharing daily progress — zeeg · 2026-10-09
- Every model looked bad in my eval — the bug was my answer key, not the models — jgarg27 · 2026-10-09
- Pi has no official Subagents, but 4 community extensions emerged; author shares extension stack order — solyarisoftware · 2026-10-09