Benchmark scores are evidence about an episode, not a job — two papers agree
Shahules786 · x · 2026-09-03
Shahules' Paper Club thread (1/6): benchmark scores are evidence about an episode, not a job — scores are often more precise than benchmark claims. Two papers make this point: CRMArena Pro (Salesforce/UIUC) and "Designing Benchmarks for Knowledge Work" (Harvard/Stanford). A model can score well on a narrow slice of work without being capable of the job around it. Later posts cover CRMArena-Pro's construction, results, limits, and the proposed reporting standard.
More from Research
- Why LLMs Count 8 People When Only 7 Are Online: A Negative-Constraint Schema for Long-Context Causality — wenger2026-12 · 2026-09-03
- Startup Mostik bridges AI models via their weights, tops ARC-AGI 3 at 1/20 the cost — nordicinst · 2026-09-03
- ByteDance's looped language models match 12B rivals at 1.4B size, with Bengio as co-author — peterjliu · 2026-09-03
- OpenAI's CoT monitorability hit: paper authors double down on 'fragile' AI safety window — GaryMarcus · 2026-09-03
- PufferLib author says retuning brings ~3x end-to-end speedup, up to 10x in some envs — yacineMTB · 2026-09-03
- Harvard Proposes Agentic Data Cracking, Cuts Unstructured QA Cost 53% on FanOutQA — Harvard · 2026-09-03