VeriHarness Paper: Disagreement Resolver Plus Consensus Challenger Lifts Agentic Verification by 6.4 Points

mikeflache · x · 2026-10-05

Researchers at Google (Caiqi Zhang, Rujun Han, et al.) released VeriHarness on arXiv, a framework for strengthening LLM-agent verification on long-horizon tasks with a fixed base model and no reference answers or rubrics at test time.

Key findings and method:

Results: Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among baselines; evidence-backed revision adds average gains of 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8 over a single rollout.

Related event: Google's VeriHarness: Disagreement May Beat Consensus in AI Evaluation(2 posts)→

Original post →

More from coding & agent

coding & agent channel →