Benchmark scores move 39 points on harness alone — a 28-min guide to building your own evals

ghumare64 · x · 2026-09-30

Rohit Ghumare argues frontier-lab benchmarks mostly exist to sell models and tools, and published a 28-minute guide on reading benchmark scores and building your own evals.

Part 1 — What a score is made of: five choices (task set, grader, harness, sampling, date) each move the number. On ARC-AGI-3, one model scored 59.3% in the standard harness but 98.4% via a provider adapter — a 39-point swing from harness alone. GPQA Diamond's 198 questions give a ±4.2-point 95% interval near 90%. SWE-Bench Pro V2 dropped 89 invalid tasks on Sept 22, 2026.

Part 2 — Build your own: failure reading, grader design, judge calibration, error bars, held-out hillclimb, agent trajectories, and production checks — works with any OpenAI-compatible endpoint and any headless-mode agent CLI, with every number reproduced on the author's laptop.

Original post →

More from Research

Research channel →