Inkling Released with Day-0 Inference Support
simonguozirui · x · 2026-07-16
Quoting content, Thinking Machines released Inkling today, along with:
- A multimodal model capable of reasoning over text, images, and audio
- Full weights open for fine-tuning on Tinker, also available on Inkling Playground
- Ecosystem already has Day-0 support, with inference stacks mentioning NVIDIA B200/B300, AMD MI350X/MI355X, native FP4 serving, unified kernel architecture, MTP, optimized attention, and heterogeneous KV cache
- The post also provides throughput numbers: 355 tok/s/user on 4×B200 + MTP
More from Infra
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22