Models systematically underestimate their own capabilities, even newest ones
repligate · x · 2026-10-08
Reshared by Anthropic's repligate: a user reports that months ago Opus 4.6, asked to guess its own specs, estimated a 250k context window and was visibly shocked upon seeing its real benchmark scores.
The user since made a habit of asking models to self-assess, and finds they are always wrong in a diminishing direction — even the newest Fables model. This blind spot in self-modeling matters for how agents plan and delegate their own work.
More from Models
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08
- User reports Haiku 5.5 is a major workflow upgrade in screenshot post — Sorcerer12345 · 2026-10-08
- OpenRouter launches Decision Model Rankings, with typesafeai leading all categories — gaganghotra_ · 2026-10-08
- OpenAI launches Intelligent UI: ChatGPT now answers with fully interactive interfaces — gdb · 2026-10-08
- Claude Haiku 5.5 Beats GPT-6 Luna on Every Benchmark but Costs 5x Past 100k Tokens — daniel_mac8 · 2026-10-08