Inkling Model Specs, Architecture Details & SGLang Support
BanghuaZ · x · 2026-07-16
Thinking Machines' open-source model Inkling boasts 975B total parameters and 41B active parameters (MoE architecture), supports up to a 1M context, and natively handles text, image, and audio reasoning.
Architecture & Inference Optimization:
- Employs a novel architecture, including ShortConv, relative positional encoding attention, and shared expert MoE.
- SGLang now offers comprehensive support: including PD disaggregation, multi-LoRA serving, HiCache, and a DFlash checkpoint trained by Modal.
- Achieves full CUDA graph optimization during the prefill phase and utilizes an MXFP8 KV cache.
More from Infra
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22