vLLM Achieves Day-0 Support for TML Inkling

vLLM Blog · rss · 2026-07-15

vLLM announced Day-0 support for Thinking Machines Lab's multimodal model, Inkling. By leveraging multi-token prediction (MTP), long-context serving, and parallel processing, vLLM optimized throughput to reach up to 380 tokens/s for a single user on NVIDIA GB200 GPUs.

Original post →

More from Infra

Infra channel →