Inkling-Small Launches on Together AI: 276B MoE with 1M Context
togethercompute · x · 2026-08-01
Together AI announced a high-throughput production path for developers to run Thinking Machine Labs' new model, Inkling-Small. Key features include:
- Architecture: A Mixture-of-Experts (MoE) transformer with 276B total parameters and 12B active per token.
- Native Multimodal: Reasons natively across text, images, and audio in a shared hidden space.
- 1M Context: Supports up to 1M tokens with variable thinking effort to balance cost and quality.
- Agentic Coding: Matches Inkling's performance on Terminal-Bench 2.1 at roughly one-fifth of the cost.
More from Models
- Surge AI Launches Off-the-Shelf Frontier Training Data, Boosting Agent Benchmarks by Double Digits — echen · 2026-08-01
- DeepSeek v4 Flash Reported to Loop and Forget Context in Coding — kwizzle · 2026-08-01
- Local Deployment on DGX Spark: Exploring Upgrades Beyond Qwen 3.5 122B — Voxandr · 2026-08-01
- 25.5 Trillion Tokens a Day: How SuperPods Power Massive RL Training — zephyr_z9 · 2026-08-01
- Benchmarking Agent Harnesses: Kimi K3 Shines, Claude Code Costs 4x More — omarsar0 · 2026-08-01
- Agentic RL Gives LLMs 'Vitality': Models Become Frantic and Engaged — teortaxesTex · 2026-08-01