Ofir Press: agents are so good it takes months to find new benchmarks they can't pass
OfirPress · x · 2026-10-02
Responding to a discussion on benchmark saturation, Princeton researcher Ofir Press says that because agents are so capable, it now takes months of hard work to find new benchmarks. As long as agents remain imperfect, though, he argues the community will keep finding benchmarks they can't solve.
Related event: Researchers Debate Benchmark Design as AI Agents Grow Stronger(3 posts)→
More from Models
- Gemini 4 Argon posts lowest hallucination rate (15%) on AA-Omniscience benchmark — import_jmr · 2026-10-02
- User says Opus 5.5 shows no token savings, hits 5-hour limit in 3-4 messages — MarsupialFirst8617 · 2026-10-02
- llama.cpp adds Decision Models, expanding local inference capabilities — paf1138 · 2026-10-02
- Early Argon impressions: dev vibes-codes with it, says it shows no signs of benchmaxxing — cgarciae88 · 2026-10-02
- NVIDIA's Kumo-Tabular Tabular Foundation Model Trends on Hugging Face — nvidia · 2026-10-02
- Claude Sonnet 5.5, Grok 4.7 and GPT-6.1 Sol go live on Runware's OpenAI-compatible endpoint — aziz4ai · 2026-10-02