Fable is likely 2-2.5T params, not 10T, as Kimi K3's 2.8T closes the gap — but model sizing is far trickier

jietang · x · 2026-09-10

Reacting to O'Neill's speculation that Fable is 2-2.5T parameters (not 10T) and that Kimi K3 — 2.8T params trained on 20-30k Blackwell-equivalents — lands close to Fable 5, jietang notes that finding the optimal model size is tricky: data volume, active parameter count, number of RL environments, and target inference cost all interact, and performance depends on many more factors, each adding variability. The unverified speculation implies Anthropic's compute, RL environments, architecture and optimizer advantages mean if Fable only edges out K3, it's almost certainly a smaller model, with GPT-5.5/5.6 smaller still.

Original post →

More from Models

Models channel →