Tinker Results Don't Prove Inkling is the Best
hero88645 · reddit · 2026-07-19
The author reviews public materials related to Thinking Machines' Tinker, questioning whether the best public results actually prove that Inkling is the optimal foundation model.
Public Benchmarks
- On AIME 2026, Inkling scores 97.1%, but higher-scoring models exist in the same table, such as GLM 5.2 at 99.2%, and Fable 5 and GPT-5.6 Sol at 99.9%.
- On HLE text-only, Inkling scores 29.7%, beating Nemotron 3 Ultra but trailing GLM 5.2, DeepSeek V4 Pro, and two Kimi models.
- On SWEBench Pro and Terminal Bench 2.1, the trend continues: Inkling beats some models but falls behind stronger open/closed-source alternatives.
- Its only standout performance is on IFBench, where it leads with 79.8%, indicating strong instruction-following capabilities.
Core Controversy
- The author points out that people often conflate two things:
- Tinker can fine-tune open-source models into highly capable specialized models.
- Inkling itself is the best foundation model to use.
- Current public evidence only supports the first point, not the second.
The Bridgewater Case
- The mentioned Bridgewater/Tinker case study used Qwen3-235B as the base, not Inkling.
- This case achieved an 84.7% accuracy across six financial document filtering tasks, with per-task inference costs roughly 13.8x lower than tested frontier models.
- Inkling did not exist when this case study was published, so it cannot be used to prove Inkling's superiority.
Conclusion
The author concludes that current materials only prove Tinker's fine-tuning/specialization pipeline is effective, but there is insufficient evidence to crown Inkling as the go-to foundation model.
More from Models
- Grok 4.5 is now free inside Cursor, the popular AI coding IDE — mark_k · 2026-07-21
- GPT often converges on the same near-miss ideas in math problems — yacineMTB · 2026-07-21
- Eno Reyes says model distillation is basically unstoppable — LangChain · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- OpenAI hackathon project stalls as Codex struggles on voice, while Claude spots the issue — ColleenMBrady · 2026-07-21
- Kimi K3 lands exactly on China’s 2-year AI capability trend line — peterwildeford · 2026-07-21