Thinking Machines Releases Open-Source Model Inkling-Small
Founded by former OpenAI CTO Mira Murati, Thinking Machines has officially released its second open-source model, Inkling-Small. Utilizing a MoE architecture with 276B total parameters and only 12B active parameters, it matches the performance of the original Inkling at a quarter of the size. It also becomes the top-scoring open-weight model on the ARC-AGI benchmark. Supporting a 1 million token context and native multimodal capabilities, the model significantly lowers deployment barriers, drawing massive community attention.
已确认
- 模型架构与规格: Inkling-Small adopts a Mixture-of-Experts (MoE) architecture with 276B total parameters and 12B active parameters per token. It supports a 1 million token context window, features variable thinking intensity, and handles native multimodal inputs from image and audio to text.
- 开源与性能表现: The model's weights are fully open under the Apache 2.0 license. Performance-wise, @ArtificialAnlys notes it rivals flagship models with only a third of the active parameters; @simonguozirui shared that it surpasses peer models and even larger versions in agentic and reasoning tasks; @keirp1 revealed that in official ARC Prize validation, it became the highest-scoring open-weight model on the ARC-AGI-1 and ARC-AGI-2 benchmarks, with a single inference cost of $0.23.
- 部署与量化支持: Various deployment options are available from the official team and community. @pcuenq confirmed its nvfp4 checkpoint runs on a single Mac Studio (MLX) using about 150GB, and can even operate across two 128GB laptops. @woosukk highlighted that the model is small enough to run on a single NVIDIA B300 accelerator. Additionally, the Unsloth team provided dynamic quantization, and the Baseten platform has simultaneously launched the model.
为什么重要
- 高性价比与低门槛: Inkling-Small demonstrates extreme parameter efficiency, achieving top-tier performance with very few active parameters and breaking the stereotype that massive models must rely on huge compute clusters. This allows developers to locally run 100B+ parameter models on consumer-grade or light enterprise hardware, greatly lowering the barrier to trialing and deploying cutting-edge AI models.
2026-07-30 ~ 2026-07-31 · 29 related posts
- Episode 1: Inkling Tops ARC-AGI Leaderboard for Open-Weight Models(2026-07-17, 2 posts)
- Episode 2: Thinking Machines Releases Open-Source Model Inkling-Small(2026-07-30, 29 posts)
Primary sources
- Unsloth Releases Inkling-Small Multimodal MoE Model — unsloth · 2026-07-30
- Inkling-Small Released: 276B Parameter MoE Model Matches Original Performance — ziqiao_ma · 2026-07-31
- Thinking Machines Releases Inkling-Small: 276B Parameters Matching Original Performance — simonguozirui · 2026-07-31
- Thinking Machines Releases Inkling-Small Model — NielsRogge · 2026-07-31
- [source] Inkling Small launches: 276B parameter model runs on a single Mac Studio with MLX — pcuenq · 2026-07-31
- thinkingmachines Releases Inkling-Small: 276B Params, 1M Context — rerri · 2026-07-31
- Thinking Machines Launches Inkling-Small: 12B Active Params Beats Larger Model — mervenoyann · 2026-07-31
- Thinking Machines Launches Inkling Small: Open-Weights Model Matches Flagship at 1/3 Size — ArtificialAnlys · 2026-07-31
- Thinking Machines Launches Inkling Small: 12B Active Params Matches Flagship Performance — ArtificialAnlys · 2026-07-31
- Inkling-Small Released: Matches Original Performance at 1/4 the Size — arshdeep · 2026-07-31
- 276B Param MoE Model Runs on Single B300, vLLM Creator Praises Inkling-Small — woosuk_k · 2026-07-31
- Running 276B Inkling Small on a Single Mac Studio via nvfp4 Quantization — pcuenq · 2026-07-31
- Inkling-Small Released: 276B Parameter MoE Model — simonguozirui · 2026-07-31
- Tinky Machines Releases Inkling-Small: 276B Parameters with Open Weights — simonguozirui · 2026-07-31
- Inkling-Small Released: Beats Larger Sibling at 1/4 the Size — simonguozirui · 2026-07-31
- Thinking Machines Releases 276B Open Model Inkling-Small — danielhanchen · 2026-07-31
- Thinking Machines Launches 276B Inkling-Small for Ultra-Fast Speech-to-Speech — MaziyarPanahi · 2026-07-31
- Inkling-Small Released: Matches Larger Model at 1/4 Size via Distillation — simonguozirui · 2026-07-31
- Inkling-Small Released: 12B Active Params Delivers Comparable Performance — clarejtbirch · 2026-07-31
- [source] Inkling-Small: New MoE Model for Image/Audio-to-Text Trends on Hugging Face — thinkingmachines · 2026-07-31
- [source] Inkling Small Tops ARC-AGI Open-Weight Leaderboard at $0.23/Task — keirp1 · 2026-07-31
- Thinking Machines Releases Inkling-Small: A 276B MoE Open Weight Model — LiTianleli · 2026-07-31
- Inkling-Small (12B Parameters) Ranks Alongside 2-3x Larger Open Models — simonguozirui · 2026-07-31
- Mira Murati's startup releases open-source Inkling-Small model — simonguozirui · 2026-07-31
5 near-duplicate retellings: simonguozirui · simonguozirui · ArtificialAnlys · simonguozirui · baseten