Anthropic: Infrastructure noise can swing agentic coding evals by over 6%

giansegato · x · 2026-08-29

Anthropic's engineering blog reveals that agentic coding benchmarks like SWE-bench are heavily influenced by infrastructure configuration. In Terminal-Bench 2.0 tests, resource allocation differences caused score swings of up to 6 percentage points, exceeding the gaps between top models. The post discusses how resource limits impact results and proposes a methodology for calibrating resources to reduce noise.

Original post →

More from coding & agent

coding & agent channel →