Google's VeriHarness shows agent agreement hides shared errors, gains 6+ points with agentic verifier

omarsar0 · x · 2026-10-04

A new Google paper challenges the standard practice of trusting agent rollouts that agree: consistent answers can hide shared errors, while disagreement often points to the correct alternative.

VeriHarness turns the same base model into an agentic verifier with two jobs:

Across five long-horizon benchmarks it posts the best selection scores among baselines tested. With evidence-backed revision it adds 6.2 points over a single rollout on Gemini 3.5 Flash and 6.4 on Claude Opus 4.8. The authors also release about 26,000 rollouts.

Original post →

More from coding & agent

coding & agent channel →