Tinker Results Don't Prove Inkling is the Best
hero88645 · reddit · 2026-07-19
The author reviews public materials related to Thinking Machines' Tinker, questioning whether the best public results actually prove that Inkling is the optimal foundation model.
Public Benchmarks
- On AIME 2026, Inkling scores 97.1%, but higher-scoring models exist in the same table, such as GLM 5.2 at 99.2%, and Fable 5 and GPT-5.6 Sol at 99.9%.
- On HLE text-only, Inkling scores 29.7%, beating Nemotron 3 Ultra but trailing GLM 5.2, DeepSeek V4 Pro, and two Kimi models.
- On SWEBench Pro and Terminal Bench 2.1, the trend continues: Inkling beats some models but falls behind stronger open/closed-source alternatives.
- Its only standout performance is on IFBench, where it leads with 79.8%, indicating strong instruction-following capabilities.
Core Controversy
- The author points out that people often conflate two things:
- Tinker can fine-tune open-source models into highly capable specialized models.
- Inkling itself is the best foundation model to use.
- Current public evidence only supports the first point, not the second.
The Bridgewater Case
- The mentioned Bridgewater/Tinker case study used Qwen3-235B as the base, not Inkling.
- This case achieved an 84.7% accuracy across six financial document filtering tasks, with per-task inference costs roughly 13.8x lower than tested frontier models.
- Inkling did not exist when this case study was published, so it cannot be used to prove Inkling's superiority.
Conclusion
The author concludes that current materials only prove Tinker's fine-tuning/specialization pipeline is effective, but there is insufficient evidence to crown Inkling as the go-to foundation model.
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11