Thinky Model Scaling and Multimodal Insights
burny_tech · x · 2026-07-16
The shared content summarizes several technical assessments of the Thinky model:
- muP becomes meaningful again: The author believes muP helps ensure returns on the scaling ladder for 1T+ parameter models.
- Low sparsity: Speculated to be caused by hardware constraints or lower MFU.
- 45T tokens: One of the highest training token counts for an open model to date. The author feels this aligns with scaling law expectations, but the model quality still appears somewhat weak, indicating insufficient gains from effective parameters.
- Unconventional architectural choices: Including convolutions in attention and the coupling of weight decay with learning rates, which the author thinks is likely tied to specific training setups.
- Native multimodal training: The author speculates the team may have learned methods from OpenAI, will prioritize this as a differentiator for open models, and guesses that multi-task training might aid coding and reasoning capabilities.
Related event: Inkling Performance and Related Architecture Speculation(7 posts)→
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11