vLLM Releases Inkling-Small Deployment Guide: Runs on Minimum 180GB VRAM
vllm_project · x · 2026-07-31
vLLM Recipes has officially released a deployment guide for Thinking Machines Lab's latest multimodal MoE model, Inkling-Small.
- Specs: 276B total parameters, 12B active per token. Natively handles text, image, and audio inputs to generate text, with up to 1M context length.
- Hardware Requirements: Offers NVFP4 and BF16 variants. The NVFP4 variant requires a minimum of 180GB aggregated VRAM (runnable with TP1 on B300/GB300 or TP2 on B200/GB200), while BF16 requires at least 600GB.
- Inference Acceleration: Supports relative-attention, short-conv serving path, tool/reasoning parsers, and MTP speculative decoding.
- Setup: Requires a vLLM nightly build; the optional vllm[audio] extra is needed for audio inputs.
More from Infra
- Open Source Project Logs Hidden LLM Serving Traps — alexcovo_eth · 2026-07-31
- Would 10k tok/s Decode Speed Unlock New LLM Use Cases? — LivingSwitch · 2026-07-31
- Inference Optimizations Yield 10x Gains, GPUs May Echo Dark Fiber Lesson — chandan1_ · 2026-07-31
- Stanford Researchers Propose Measuring AI Efficiency by 'Intelligence Per Watt' — StanfordAILab · 2026-07-31
- New Podcast Episode Asks: Do Data Centers Consume a Lot of Water? — AndyMasley · 2026-07-31
- DeepSeek Plans 1GW AI Data Center in Inner Mongolia, Anthropic Targets 9-10GW by 2027 — zephyr_z9 · 2026-07-31