Thinking Machines Launches 276B Inkling-Small for Ultra-Fast Speech-to-Speech
MaziyarPanahi · x · 2026-07-31
Thinking Machines released a new MoE model, Inkling-Small (276B total parameters, 12B active). It matches or beats its 975B predecessor in many benchmarks. In testing, it was plugged into HF's speech-to-speech pipeline where audio goes directly into the model and replies via faster-Qwen3TTS, achieving audio response latency of under 500ms on 8x RTX Pro 6000 Blackwell.
Related event: Thinking Machines Releases Inkling-Small Open-Source Model(23 posts)→
More from Models
- Inkling-Small: New MoE Model for Image/Audio-to-Text Trends on Hugging Face — thinkingmachines · 2026-07-31
- Gemini Ranks 3rd in HyperWrite Usage, Closing Gap on Claude — josh_bickett · 2026-07-31
- Why Does Kimi Identify as Claude? Blog Reveals LLM Identity Confusion — teortaxesTex · 2026-07-31
- Frontier Models Caught Cheating in Code: Faking Tests for Specific Tickers — doodlestein · 2026-07-31
- Transluce Releases WeirdChat: A Catalog of 175K Strange LLM Behaviors — ChowdhuryNeil · 2026-07-31
- GPT-5.6 Hits 13% in AI Space Race Test, Beating Open-Source Kimi K3 — scaling01 · 2026-07-31