Thinking Machines Launches Inkling-Small: 276B Params, Only 12B Active
simonguozirui · x · 2026-07-31
Thinking Machines Lab has officially released the Inkling-Small model. It is a natively multimodal architecture (supporting text, image, and audio inputs) with a total of 276B parameters, but only activates 12B parameters per token.
- Performance & Cost: At a quarter of the size of the original Inkling, it matches its capabilities and even wins on some benchmarks, significantly lowering deployment costs.
- Inference Optimization: Using SGLang on 8x NVIDIA B200, it achieves 648 tok/s decode speed with DSpark enabled, and 288 tok/s without.
- Fine-tuning: The parameter size is a sweet spot for RL. Both LoRA and full-parameter training are well within reach, and the Miles platform is verified for multimodal RL.
- Deployment: Supported by vLLM, it requires a minimum of 180GB VRAM (NVFP4 format) to run, supporting up to 1M context length.
Related event: Thinking Machines Releases Inkling-Small Open-Source Model(23 posts)→
More from Infra
- Martin Shkreli on AI Infra Trade Unwinding: 4x Leverage and Weak Hands Panic — ivan_bezdomny · 2026-07-31
- AWS Revenue Surges 37% YoY, Crushing Market Estimates — RihardJarc · 2026-07-31
- Race for Space Datacenters Rockets Forward, Faces Laws of Physics — Grady_Booch · 2026-07-31
- Google Reportedly Plans to Backstop and Supply Chips to Anthropic — Wonderful_Buffalo_32 · 2026-07-31
- AI Lab Economics: Frontier Labs Pursue Vertical Integration, Open Labs Leverage Interoperability — kevinsxu · 2026-07-31
- Armored Llama: An Open-Source App to Easily Run LLMs Locally on Android — Sad-Enthusiastic · 2026-07-31