Fable is likely 2-2.5T params, not 10T, as Kimi K3's 2.8T closes the gap — but model sizing is far trickier
jietang · x · 2026-09-10
Reacting to O'Neill's speculation that Fable is 2-2.5T parameters (not 10T) and that Kimi K3 — 2.8T params trained on 20-30k Blackwell-equivalents — lands close to Fable 5, jietang notes that finding the optimal model size is tricky: data volume, active parameter count, number of RL environments, and target inference cost all interact, and performance depends on many more factors, each adding variability. The unverified speculation implies Anthropic's compute, RL environments, architecture and optimizer advantages mean if Fable only edges out K3, it's almost certainly a smaller model, with GPT-5.5/5.6 smaller still.
More from Models
- Hobbyist's 348M model hits 99.4% on GPT-3 arithmetic tasks, beats 175B giant — nkthebass · 2026-09-10
- Liang Wenfeng Stars in a Promo Video for DeepSeek V4.1 Flash, Complete With a Catgirl — teortaxesTex · 2026-09-10
- Naval: Frontier labs' flywheel is distilling data from the smartest users of leading models — naval · 2026-09-10
- DeepSeek V4.1 Flash Hits the App: One Model Replaces Three Modes, Claims It Beats V4 Pro — teortaxesTex · 2026-09-10
- Kimi K3 scores 60% higher than Fable 5.1 on Harvey's hard autonomous legal tasks — togethercompute · 2026-09-10
- OpenAI internal model solves Navier-Stokes in 88 hours using ~10,000 coordinating agents — TheMoonMidas · 2026-09-10