giffmana: current models are great at precision but really bad at recall

giffmana · x · 2026-10-02

Replying to Ofir Press on benchmark saturation, giffmana reframes the issue: such benchmarks test recall, and in his hands-on experience current-generation models are really good at precision but really not good at recall.

Related event: Researchers Debate Benchmark Design as AI Agents Grow Stronger(3 posts)→

Original post →

More from Models

Models channel →