Eval author warns against over-indexing on benchmarks; Anthropic finds infra noise can swing scores 6 points

giansegato · x · 2026-09-29

giansegato confirms he ran the evals multiple times (details in the system card PDF) and cautions that industry benchmarks carry confounders that are hard to fully control—better to just try the model. His reply links to Anthropic's engineering post "Quantifying infrastructure noise in agentic coding evals," showing infrastructure configuration alone can shift agentic coding benchmark scores by several percentage points—on Terminal-Bench 2.0, the gap between best and worst-resourced setups was 6 percentage points (p < 0.01), sometimes exceeding the margin separating top leaderboard models.

Original post →

More from Research

Research channel →