Initial probing of Opus suggests its time estimates are hallucinated from human training samples
StefanoGogioso · x · 2026-10-01
Responding to the debate over Claude's "2-3 days" estimates, Stefano Gogioso ran a preliminary investigation into Opus 5.5.
He found the numbers appear to be inferred or hallucinated from human samples in the training data, with little to no attempt at self-reflection — unless the context has already narrowed to an agent persona.
His conclusion: hypothesis 3 (statistical imitation from training data) is most likely.
More from Models
- Rumor: Google's Gemini 4 Argon spotted ahead of launch, details still unconfirmed — koltregaskes · 2026-10-01
- Google is back: Gemini 4 Argon beats Astra and Opus 5.5 across the board — Yuchenj_UW · 2026-10-01
- Argon hits 55% on FrontierSWE v2 long-horizon coding, vs ~20% for Gemini 3.x — xennygrimmato_ · 2026-10-01
- Full benchmark results for Gemini 4 Argon with high reasoning released — ArtificialAnlys · 2026-10-01
- Gemini 4 Argon tops AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 — ArtificialAnlys · 2026-10-01
- GPT-6.1 Sol cached input at $0.10/M: the line that sets your agent bill — victor_explore · 2026-10-01