giffmana: current models are great at precision but really bad at recall
giffmana · x · 2026-10-02
Replying to Ofir Press on benchmark saturation, giffmana reframes the issue: such benchmarks test recall, and in his hands-on experience current-generation models are really good at precision but really not good at recall.
Related event: Researchers Debate Benchmark Design as AI Agents Grow Stronger(3 posts)→
More from Models
- Looped Transformers become the latest research meta, says Kye Gomez — KyeGomezB · 2026-10-02
- Which $400+ AI agent is actually best? One idea: make them compete in a city — nick_linck · 2026-10-02
- fastino launches GLiDE, an on-demand reasoning decision model that tops the Decision Index by 6.9 points — kalyan_kpl · 2026-10-02
- Anthropic Reveals How User Feedback Reshaped Fable 5.1 After 'Hard to Talk To' Sonnet 5 — every · 2026-10-02
- 1,565-email benchmark: cheapest Perplexity/Cloudflare decision model is also most accurate — michellechen · 2026-10-02
- Two 96GB Huawei Ascend cards run Qwen locally: from 1 tok/s to 30 tok/s — matteiuspi · 2026-10-02