Eval author warns against over-indexing on benchmarks; Anthropic finds infra noise can swing scores 6 points
giansegato · x · 2026-09-29
giansegato confirms he ran the evals multiple times (details in the system card PDF) and cautions that industry benchmarks carry confounders that are hard to fully control—better to just try the model. His reply links to Anthropic's engineering post "Quantifying infrastructure noise in agentic coding evals," showing infrastructure configuration alone can shift agentic coding benchmark scores by several percentage points—on Terminal-Bench 2.0, the gap between best and worst-resourced setups was 6 percentage points (p < 0.01), sometimes exceeding the margin separating top leaderboard models.
More from Research
- BLUE from Google internship lands NeurIPS: LLM-written user profiles boost recommendations — shangbinfeng · 2026-09-29
- Collatz twist: Krasikov-Lagarias-style X^0.84 bounds apply to any root, and equally to 3x-1 — AlexKontorovich · 2026-09-29
- Emergence debunked? Researcher argues emergent LLM abilities are labels in your imagination — gerardsans · 2026-09-29
- MHCflurry 2.3 ships with multi-GPU speedups and first new model weights in ages — iskander · 2026-09-29
- 2,000+ navigation tasks across 133 environments test how far LLMs are from zero-shot robot control — PaulYacoubian · 2026-09-29
- Microsoft Research Asia Singapore marks year one: healthcare AI deployment and 75+ local projects — Microsoft Research · 2026-09-29