Thinky Model Scaling and Multimodal Insights
burny_tech · x · 2026-07-16
The shared content summarizes several technical assessments of the Thinky model:
- muP becomes meaningful again: The author believes muP helps ensure returns on the scaling ladder for 1T+ parameter models.
- Low sparsity: Speculated to be caused by hardware constraints or lower MFU.
- 45T tokens: One of the highest training token counts for an open model to date. The author feels this aligns with scaling law expectations, but the model quality still appears somewhat weak, indicating insufficient gains from effective parameters.
- Unconventional architectural choices: Including convolutions in attention and the coupling of weight decay with learning rates, which the author thinks is likely tied to specific training setups.
- Native multimodal training: The author speculates the team may have learned methods from OpenAI, will prioritize this as a differentiator for open models, and guesses that multi-task training might aid coding and reasoning capabilities.
Related event: Inkling Performance and Related Architecture Speculation(7 posts)→
More from Models
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Ben’s Bites roundup highlights Kimi K3, Fable 5, Cursor costs and self-driving companies — Ben's Bites · 2026-07-21