Anthropic's Minor .1 Releases Show Benchmark Jumps That Make No Sense
thursdai_pod · x · 2026-09-07
The Thursdai podcast flags that Anthropic's Fable 5.1 and Mythos 5.1 releases, despite being minor .1 version bumps, show implausibly large benchmark jumps — including on Terminal-Bench 4, which didn't even exist last week. @altryne breaks down why these numbers look off.
More from Models
- GPT-6 Astra generates animated black hole scene with custom WebGL/GLSL shaders — omarsar0 · 2026-09-07
- 146 model variants tested on creative writing: Opus, Fable, Kimi K3 top the board — zainhas · 2026-09-07
- Claude's Writing Tics: Personified Subjects and Arguing by Negation, Dissected — Sauers_ · 2026-09-07
- GPT-6 Astra on Low Beats GPT-5.6 Sol on High, Devs Advise Lower Effort — reach_vb · 2026-09-07
- After 115K videos in prod, engineer shares Gemini video-understanding gotchas and hacks — TheMoonMidas · 2026-09-07
- TheZvi: Astra fails EditorBench for the same reasons as Sol — TheZvi · 2026-09-07