Google's VeriHarness: Disagreement May Beat Consensus in AI Evaluation
Google researchers introduced VeriHarness on arXiv, challenging the assumption that majority agreement among agent rollouts means correctness. By resolving disagreements and challenging consensus, the framework boosts long-horizon task verification to 6.4 without reference answers.
2026-10-04 ~ 2026-10-05 · 2 related posts