ThinkingMachines' Distilled Small Model Beats Large Model in Coding and Reasoning
量子位 · wechat · 2026-07-31
ThinkingMachines released the Inkling-Small model (276B total parameters, 12B active). Surprisingly, this distilled model outperforms its larger teacher model on certain metrics: scoring 31.6% on HLE (vs. 29.7% for the large model) and breaking 80% on SWEBench Verified (vs. 77.6%).
This isn't just cherry-picking benchmarks; at equivalent compute costs, the small model's test-time compute curve consistently stays above the large model's. The official recipe for this success involves leveraging the time gap to adjust pre-training data ratios, applying online policy distillation using the large model as a teacher, and finishing with two extra weeks of RL specifically for agentic coding.
However, the announcement also notes the small model's limits: on hard knowledge and factual tasks that rely on parameter scale (like SimpleQA Verified), the large model still holds a decisive advantage.
Related event: Mira Murati's Startup Releases Open-Source MoE Model Inkling-Small(41 posts)→
More from Models
- Google DeepMind Unveils Gemini Robotics 2 with Advanced Dexterity — ___Mufasaa · 2026-07-31
- Deep20Bench tests LLM strategy via 'Twenty Questions': Opus 5 and Kimi K3 lead the pack — wauwau0977 · 2026-07-31
- DeepSeek-V4-Flash surpasses V4-Pro-Preview in latest benchmarks — Outside-Risk-8912 · 2026-07-31
- DeepSeek-V4-Flash-0731 far surpasses DeepSeek-V4-Pro-Preview in benchmarks — SnooBunnies8392 · 2026-07-31
- DeepSeek's Price War Sparks Hype for DeepSeek 4 Pro — kimmonismus · 2026-07-31
- Codex on ChatGPT Business Costs $300/Day Extra, Devs Prefer Claude — paul_cal · 2026-07-31