Inkling: 1T Open-Source Multimodal Model
SonglinYang4 · x · 2026-07-16
vLLM shared an introduction to Thinking Machines' Inkling: a 1T parameter open-weights model supporting native multimodal input across text, image, and audio.
Key Details
- Maximum context length of up to 1 million tokens
- New architecture: relative attention, short convolutions, MoE expert sinks
- Equipped with 8 MTP heads for speculative decoding
- Supports NVFP4 and BF16 checkpoints
- Optimized for NVIDIA Blackwell and Hopper
- Achieves 380 tok/s/user on a 4× GB200 setup when combined with MTP
The official announcement confirms that full weights are open, available for fine-tuning on Tinker, and ready to test in the Inkling Playground.
More from Multimodal
- Gemini Flash 3.8 image-to-SVG test sparks claim SVG may replace image models in 18 months — Kyrannio · 2026-09-03
- Higgsfield's new Genjutsu motion-copy tool impresses: better than Kling motion control? — rheylew · 2026-09-03
- Team claims h3 max is the undisputed #1 frontier video model across benchmarks — isidentical · 2026-09-03
- Fable 5.1 makes three.js sites: faster and sharper, but taste still matters — repligate · 2026-09-03
- Possible open-source MiniMax H3 Max weights appear on Hugging Face, real-time on 8x B200 — BassNet · 2026-09-03
- Point-and-click adventure built by chaining Nano Banana 2, MiniMax H3 Max, SAM 3 and GPT-5.6 — yshan2u · 2026-09-03