FULL STORY
Inkling: From Open-Weight Debut to Performance Dispute
Thinking Machines Lab launched its first open-weight multimodal model, Inkling, with a simultaneous debut on Databricks. The release quickly led to community architecture sleuthing, but early hands-on tests produced disappointing feedback and a polarized reception.
2026-07-16 ~ 2026-07-17 · 3 episodes · 178 posts
Episode 1 · Thinking Machines launches Inkling, its first open-weight multimodal model (2026-07-16, 169 posts)
Thinking Machines has released Inkling, its first open-weight model and the first entry in a broader model family. The launch drew attention because multiple posts described it as a 975B-total, 41B-active multimodal MoE with a 1M context window, and because the company paired the announcement with Tinker fine-tuning, a Playground, and immediate ecosystem availability.
Key details
The official account said Inkling can perform efficient reasoning across text, images, and audio, and that full weights will be opened. Multiple posts relaying the model card described it as roughly a 1T-parameter model under an Apache-2 license, with support for a 1M-token context window; they also said it was trained from scratch on 48T tokens spanning text, images, and audio. The company later added that Inkling is only the start of a series: a lighter Inkling-Small with a similar training recipe has already been previewed, and more complete weights will continue to be released.
Launch and ecosystem
Beyond Tinker and the Playground, a Databricks announcement relayed by @Yuchenj_UW said Inkling was available on Databricks as a Day 0 launch partner. The company also highlighted that thinkingmachines/Inkling appeared on Hugging Face trending, where it was labeled as an image-text-to-text pipeline and also showed audio-text-to-text support.
Reactions
On positioning, @natolambert said the model card and benchmark information suggest Inkling is clearly stronger than Nemotron Ultra. @deedydas described it as the strongest open-weight model outside China. @bindureddy was more cautious, saying the main significance is that the U.S. open-weight camp finally has a model serious enough to benchmark closely.
- Inkling Evaluation and Positioning Revealed — natolambert · 2026-07-16
- Inkling: A 1T Parameter Open-Source Multimodal Model — natolambert · 2026-07-16
- Inkling Safety Eval Shows No Significant Risk Increase — natolambert · 2026-07-16
- Inkling Released: A Cross-Modal Reasoning Model — thinkymachines · 2026-07-16
- Inkling Released: Multimodal with Strong Audio Capabilities — thinkymachines · 2026-07-16
- Inkling Highlights Cost and Latency Advantages — thinkymachines · 2026-07-16
- Inkling Family Will Include a Small Version — thinkymachines · 2026-07-16
- Inkling to Receive a Lightweight Version — thinkymachines · 2026-07-16
- Inkling Released with Open Weights — TheZachMueller · 2026-07-16
- Thinking Machines Releases First Open-Weight Model — WhyLifeIs4 · 2026-07-16
- Thinking Machines Releases First Open-Weight Model — WhyLifeIs4 · 2026-07-16
- Inkling: Open-Weight Multimodal Model — markjeffrey · 2026-07-16
- Thinking Machines Releases Inkling — daniel_mac8 · 2026-07-16
- Inkling Released with Day-0 Inference Support — simonguozirui · 2026-07-16
- New Open-Source Multimodal Model Announced — Xianbao_QIAN · 2026-07-16
- Inkling Open-Weight Multimodal Model Released — TheZachMueller · 2026-07-16
- Inkling Model Architecture Closely Mirrors DeepSeek V3 — a_karvonen · 2026-07-16
- Inkling Model Specs, Architecture Details & SGLang Support — BanghuaZ · 2026-07-16
- Thinking Machines Releases Open-Source Multimodal Model Inkling — clarejtbirch · 2026-07-16
- Thinking Machines Engineer Retrospective on Open-Source Model Development — jonlachman · 2026-07-16
Episode 2 · Inkling Performance and Related Architecture Speculation (2026-07-16, 7 posts)
@nrehiew_ posted a rapid series of analyses on July 16 around the newly released Inkling model, starting with benchmark behavior and product positioning, then extending into architecture and training-signal discussion around a related frontier-style design. The thread drew interest for two reasons: why a model said to be about 3.5x smaller can still post unusually strong results, and whether very large MoE systems are making new trade-offs in attention, positional encoding, and training setup.
Inkling performance and positioning
In @nrehiew_’s reading, Inkling is aimed more at general-purpose reasoning than at being a pure coding agent, which he said fits its dependence on the Tinker platform and a strategy he compared to Cohere. He also pointed to tasks such as TerminalBench and SimpleQA as useful lenses for seeing what parts of model capability are doing the work in these results.
Architecture details under discussion
In a separate breakdown of a suspected frontier model design, @nrehiew_ said the attention stack uses a sliding window and basic 1D convolutions on KV and residual paths. He also highlighted the absence of RoPE, replaced instead by relative position bias collected from distance information. On the MoE side, he summarized figures of about 975B total parameters, 41B active parameters, and roughly 45T multimodal training tokens, adding that both active parameter count and token scale appear larger than Dsv3.
What remains unconfirmed
@nrehiew_ further speculated that a mechanism described as adjusting cost by token might effectively be changing the length penalty according to “thinking intensity,” which could help explain a more compressed CoT style. He explicitly framed this as inference from observed results rather than an official confirmation. A repost shared by @burny_tech also relayed community-side views that muP may matter again at the 1T+ scale, and that relatively low sparsity could reflect hardware constraints or lower expert parallelism; those points likewise were not official disclosures.
- Thinky Model Scaling and Multimodal Insights — burny_tech · 2026-07-16
- Further Details on MoE Configuration — nrehiew_ · 2026-07-16
- Suspected Frontier Model Architecture: No RoPE and Hybrid Attention — nrehiew_ · 2026-07-16
- Deep Dive into a Model's Attention and MoE Details — nrehiew_ · 2026-07-16
- Adjusting Cost by Reasoning Intensity? — nrehiew_ · 2026-07-16
- Can Small Models Be Incredibly Strong Too? — nrehiew_ · 2026-07-16
- Analyzing Inkling: General Reasoning in an Ultra-Compact Size — nrehiew_ · 2026-07-16
Episode 3 · New Open-Source Model Inkling Falls Short in Tests (2026-07-16, 2 posts)
Despite recent hype, the new open-weight model Inkling has shown instability and rough performance in user tests. Its chain of thought tends to diverge even on simple requests, lagging noticeably behind frontier open-source models.
- Inkling Falls Short of Frontier Open-Source Models — emollick · 2026-07-16
- Inkling Hands-On Test Fails to Impress — emollick · 2026-07-16