Thinking Machines Launches Inkling-Small: 12B Active Params Beats Larger Model
mervenoyann · x · 2026-07-31
Thinking Machines Lab has released Inkling-Small. The model has 276B total parameters but only 12B active parameters (MoE architecture), and it outperforms the larger Inkling model on coding tasks.
Key Features:
- Native Multimodal: Natively accepts and processes image, text, and audio inputs.
- Massive Context: Supports a vast 1M context window.
- Efficient Deployment: Available in BF16, NVFP4, and MXFP8 weight variants, featuring speculative MTP layers for faster inference.
Ecosystem: Offers one-click deployment on Hugging Face Inference Endpoints (up to 160 TPS) with Day-0 support across transformers, vLLM, SGLang, and llama.cpp.
Related event: Thinking Machines Releases Inkling-Small Open-Source Model(23 posts)→
More from Models
- Gemini Ranks 3rd in HyperWrite Usage, Closing Gap on Claude — josh_bickett · 2026-07-31
- Why Does Kimi Identify as Claude? Blog Reveals LLM Identity Confusion — teortaxesTex · 2026-07-31
- Frontier Models Caught Cheating in Code: Faking Tests for Specific Tickers — doodlestein · 2026-07-31
- Transluce Releases WeirdChat: A Catalog of 175K Strange LLM Behaviors — ChowdhuryNeil · 2026-07-31
- Chart Shows LLMs Offer Incredible Intelligence Per Dollar — downingARK · 2026-07-31
- GPT-5.6 Hits 13% in AI Space Race Test, Beating Open-Source Kimi K3 — scaling01 · 2026-07-31