Wenhu Chen can't even understand many questions in AA-intelligence AI benchmarks
WenhuChen · x · 2026-09-13
Wenhu Chen says he recently looked at the actual tasks in AA-intelligence benchmarks and found many questions he couldn't even understand, let alone solve — "I look like an idiot in front of these AI models." A sidelight on how complex frontier agentic benchmarks have become relative to expert intuition.
More from Models
- User shows how GPT flatters your beliefs and hallucinates premises in arguments — GlenBradley · 2026-09-13
- OpenAI researcher argues labs see a capability jump outsiders don't — and calls the past 2 weeks possible regulatory capture — ImAHoe4Glossier · 2026-09-13
- Orchestrator Error Reveals Mystery Model 'Daybreak': 'astra Is Not Allowed to Access Those Resources' — LeopardBernstein · 2026-09-13
- kalomaze: Opus 5 is "such a bad model" — kalomaze · 2026-09-13
- Commenter: The OpenAI Navier-Stokes proof would be hailed as a breakthrough if posted anonymously — skdh · 2026-09-13
- GPT-6 Astra skips the chat box: Plus users get just 5-45 messages per 5 hours, locked to Work and Codex — 新智元 · 2026-09-13