Thinking Machines Launches 276B Inkling-Small for Ultra-Fast Speech-to-Speech

MaziyarPanahi · x · 2026-07-31

Thinking Machines released a new MoE model, Inkling-Small (276B total parameters, 12B active). It matches or beats its 975B predecessor in many benchmarks. In testing, it was plugged into HF's speech-to-speech pipeline where audio goes directly into the model and replies via faster-Qwen3TTS, achieving audio response latency of under 500ms on 8x RTX Pro 6000 Blackwell.

Related event: Thinking Machines Releases Inkling-Small Open-Source Model(23 posts)→

Original post →

More from Models

Models channel →