Inkling-Small Launches: 276B Params, 12B Active at One-Fifth the Cost

togethercompute · x · 2026-07-31

Thinking Machines Lab has released Inkling-Small, a new open-weight multimodal model now available on Together AI.

Built on a Mixture-of-Experts (MoE) architecture, the model has 276B total parameters with only 12B active per token. It natively reasons across text, image, and audio, supporting a context window of up to 1M tokens. The model is reported to match the performance of the larger Inkling model on Terminal-Bench 2.1 at roughly one-fifth of the cost, optimized for coding, agents, and general multimodal workloads.

Related event: Thinking Machines Unveils Inkling-Small(40 posts)→

Original post →

More from Models

Models channel →