Anthropic: Infrastructure noise can swing agentic coding evals by over 6%
giansegato · x · 2026-08-29
Anthropic's engineering blog reveals that agentic coding benchmarks like SWE-bench are heavily influenced by infrastructure configuration. In Terminal-Bench 2.0 tests, resource allocation differences caused score swings of up to 6 percentage points, exceeding the gaps between top models. The post discusses how resource limits impact results and proposes a methodology for calibrating resources to reduce noise.
More from coding & agent
- Fix Cursor Compatibility: Use GPT Models via API — VraserX · 2026-08-29
- MCP Usage Explodes: Vercel Tool Calls Up 564% in 3 Months — jasonkneen · 2026-08-29
- From Chatbots to Doing Work: X17z Demonstrates Terminal Operators — Scobleizer · 2026-08-29
- Agents on Omarchy create tools for themselves, including a task workbench — BLUECOW009 · 2026-08-29
- 'Abundant Constraints Beat Abundant Implementation': An Essay on Directing AI Capability — aishashok14 · 2026-08-29
- fbtee 4.0 released, fully rewritten in Rust with Oxc — cnakazawa · 2026-08-29