Benchmark scores move 39 points on harness alone — a 28-min guide to building your own evals
ghumare64 · x · 2026-09-30
Rohit Ghumare argues frontier-lab benchmarks mostly exist to sell models and tools, and published a 28-minute guide on reading benchmark scores and building your own evals.
Part 1 — What a score is made of: five choices (task set, grader, harness, sampling, date) each move the number. On ARC-AGI-3, one model scored 59.3% in the standard harness but 98.4% via a provider adapter — a 39-point swing from harness alone. GPQA Diamond's 198 questions give a ±4.2-point 95% interval near 90%. SWE-Bench Pro V2 dropped 89 invalid tasks on Sept 22, 2026.
Part 2 — Build your own: failure reading, grader design, judge calibration, error bars, held-out hillclimb, agent trajectories, and production checks — works with any OpenAI-compatible endpoint and any headless-mode agent CLI, with every number reproduced on the author's laptop.
More from Research
- Diffusion Models Tutorial Accepted to NeurIPS 2026 Alongside 7 Paper Acceptances — mittu1204 · 2026-09-30
- AMB3R-SLAM: Kilometer-Scale Real-Time SLAM on One Consumer GPU, Cutting ATE by 70% — rsasaki0109 · 2026-09-30
- Prefix-Reuse FLOPs: new metric exposes hidden cost of arbitrary context edits in LLM serving — RulinShao · 2026-09-30
- Tencent Hunyuan releases ExplorationBench to measure how AI systems explore — TencentHunyuan · 2026-09-30
- Fully open MolmoAct 2 tops independent robotics benchmark LIBERO-MAX on dynamic robustness — DJiafei · 2026-09-30
- CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet — CompVis · 2026-09-30