Thinking Machines Unveils Inkling-Small Model
Thinking Machines Lab has officially released the Inkling-Small model. Built on a Mixture-of-Experts (MoE) architecture, this natively multimodal model features a total of 276B parameters while activating only 12B per token and supports a context window of up to 1M. Its weights are now fully open-sourced, with both the official team and the community highlighting its significantly lowered deployment barriers alongside outstanding performance.
已确认
- 架构与参数: The model has a total parameter count of 276B with 12B active parameters, making it a quarter of the size of the original Inkling. It supports text, image, and audio inputs, featuring a natively multimodal architecture.
- 上下文能力: Supports a massive context window of up to 1M.
- 开放权重与部署: Model weights are fully open. The official team has provided an NVFP4 quantized version, with synchronized support from the Unsloth team. Author @pcuenq noted that its NVFP4 checkpoint can run on a single Mac Studio (MLX), occupying about 150GB of memory, or be distributed across two 128GB laptops.
为什么重要
- 性能表现: Multiple authors noted that despite its significantly reduced size, Inkling-Small's performance closely matches the original Inkling. Author @mervenoyann emphasized that its coding capabilities even surpass those of the larger Inkling version, demonstrating exceptional parameter efficiency.
- 降低硬件门槛: While ultra-large models typically require massive compute clusters, Inkling-Small leverages extremely low active parameters and quantization techniques to run a 276B-scale model smoothly on consumer or prosumer hardware (like the Mac Studio). This is highly significant for practical applications within the open-source community and among local developers.
2026-07-31 ~ 2026-07-31 · 22 related posts
Primary sources
- Inkling-Small Released: 276B Parameter MoE Model Matches Original Performance — ziqiao_ma · 2026-07-31
- Thinking Machines Releases Inkling-Small: 276B Parameters Matching Original Performance — simonguozirui · 2026-07-31
- Thinking Machines Releases Inkling-Small Model — NielsRogge · 2026-07-31
- [source] Inkling Small launches: 276B parameter model runs on a single Mac Studio with MLX — pcuenq · 2026-07-31
- [source] thinkingmachines Releases Inkling-Small: 276B Params, 1M Context — rerri · 2026-07-31
- Thinking Machines Launches Inkling-Small: 12B Active Params Beats Larger Model — mervenoyann · 2026-07-31
- Thinking Machines Launches Inkling Small: Open-Weights Model Matches Flagship at 1/3 Size — ArtificialAnlys · 2026-07-31
- Thinking Machines Launches Inkling Small: 12B Active Params Matches Flagship Performance — ArtificialAnlys · 2026-07-31
- Inkling-Small Released: Matches Original Performance at 1/4 the Size — arshdeep · 2026-07-31
- [source] 276B Param MoE Model Runs on Single B300, vLLM Creator Praises Inkling-Small — woosuk_k · 2026-07-31
- Running 276B Inkling Small on a Single Mac Studio via nvfp4 Quantization — pcuenq · 2026-07-31
- Inkling-Small Released: 276B Parameter MoE Model — simonguozirui · 2026-07-31
- Tinky Machines Releases Inkling-Small: 276B Parameters with Open Weights — simonguozirui · 2026-07-31
- Inkling-Small Released: Beats Larger Sibling at 1/4 the Size — simonguozirui · 2026-07-31
- Thinking Machines Releases 276B Open Model Inkling-Small — danielhanchen · 2026-07-31
- Thinking Machines Launches 276B Inkling-Small for Ultra-Fast Speech-to-Speech — MaziyarPanahi · 2026-07-31
- Inkling-Small Released: Matches Larger Model at 1/4 Size via Distillation — simonguozirui · 2026-07-31
5 near-duplicate retellings: simonguozirui · simonguozirui · ArtificialAnlys · simonguozirui · baseten