Anthropic: Infrastructure Config Alone Swings Agentic Coding Benchmarks by 6 Points
giansegato · x · 2026-09-29
An Anthropic engineering post shows that infrastructure configuration alone can significantly perturb agentic coding eval results—sometimes exceeding the leaderboard gap between top models.
- In internal experiments, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points (p < 0.01), while top leaderboard spots are often separated by just a few points.
- Unlike static benchmarks, agentic evals give models a full environment to write code, run tests, install dependencies, and iterate—the runtime becomes an integral part of problem-solving, so two agents with different resource budgets and time limits aren't taking the same test.
- Terminal-Bench 2.0 now specifies recommended CPU/RAM per task, but specifying resources isn't the same as enforcing them consistently, and enforcement methodology itself can change what the benchmark actually measures.
- The finding emerged when Anthropic's GKE-cluster scores didn't match the official leaderboard during calibration.
The poster adds a caveat: industry benchmarks carry many hard-to-control confounders—don't overindex on them; just try the model.
More from coding & agent
- exe.dev kills per-seat pricing, shifts to compute pools; individual plan drops to $15/mo — davidcrawshaw · 2026-09-29
- Burkov: Asking LLMs for a plan yields over-engineering — "fix it" works better for novices — burkov · 2026-09-29
- We killed lines-of-code metrics, so why are we measuring AI productivity by merged PRs? — brandon_galang · 2026-09-29
- Every's new skill "Is This Anything" turns your unfinished AI experiments into lessons — every · 2026-09-29
- Agent Builder's Lesson: One Smart Model Surrounded by Dumb, Reliable Deterministic Parts — Miserable_Donut8718 · 2026-09-29
- LlamaIndex explores on-the-fly model routing for document parsing tasks — llama_index · 2026-09-29