Inkling-Small Launches: 276B Params, 12B Active at One-Fifth the Cost
togethercompute · x · 2026-07-31
Thinking Machines Lab has released Inkling-Small, a new open-weight multimodal model now available on Together AI.
Built on a Mixture-of-Experts (MoE) architecture, the model has 276B total parameters with only 12B active per token. It natively reasons across text, image, and audio, supporting a context window of up to 1M tokens. The model is reported to match the performance of the larger Inkling model on Terminal-Bench 2.1 at roughly one-fifth of the cost, optimized for coding, agents, and general multimodal workloads.
Related event: Thinking Machines Unveils Inkling-Small(40 posts)→
More from Models
- Google Responds to AI Misinformation Concerns: Gemini Images Embed SynthID Watermarks — henkvaness · 2026-07-31
- Users report OpenAI's o1-pro model got slower but smarter — teortaxesTex · 2026-07-31
- DeepSeek V4 Could Continue Pretraining with MOPD Reusing Domain Experts — teortaxesTex · 2026-07-31
- MiniMax H3 Video Model Enters Chatbot Arena, Open Weights Coming Soon — arena · 2026-07-31
- 2-bit Quantized Qwen 35B Evaluated on Terminal-Bench for Agentic Coding — DavidBennett__ · 2026-07-31
- OpenAI Slashes GPT-5.6 Prices by 80%, Inference Cost Drops 2000x Annually — Latent Space · 2026-07-31