Google's VeriHarness shows agent agreement hides shared errors, gains 6+ points with agentic verifier
omarsar0 · x · 2026-10-04
A new Google paper challenges the standard practice of trusting agent rollouts that agree: consistent answers can hide shared errors, while disagreement often points to the correct alternative.
VeriHarness turns the same base model into an agentic verifier with two jobs:
- Resolve claims where rollouts disagree by checking workspace evidence
- Challenge claims every rollout agrees on, hunting for requirements they all missed
Across five long-horizon benchmarks it posts the best selection scores among baselines tested. With evidence-backed revision it adds 6.2 points over a single rollout on Gemini 3.5 Flash and 6.4 on Claude Opus 4.8. The authors also release about 26,000 rollouts.
More from coding & agent
- Hosted chat UI for OpenAI's Agents API pitched: plug in agents, no FE code — DarasStayHome · 2026-10-04
- Alloyqa lets coding agents work in the cloud and hand back a PR to review — AdventurousAge8767 · 2026-10-04
- OpenAI releases 34-page whitepaper on how it builds AI agents — mdancho84 · 2026-10-04
- Meta engineer: an AI session fixes top JS errors daily and pings owners — Vjeux · 2026-10-04
- 19-year-old builds Telegram AI agent: 'It's just a wrapper, that's all' — PowerOk7047 · 2026-10-04
- Monitoring Claude/Codex all day is frying my dopamine system, dev says — ethanCaballero · 2026-10-04