276B Param MoE Model Runs on Single B300, vLLM Creator Praises Inkling-Small
woosuk_k · x · 2026-07-31
Woosuk Kwon, creator of vLLM, reshared and praised the newly released Inkling-Small by Thinkymachines, noting that it is small enough to run on a single NVIDIA B300 GPU and might be the best model in its size class.
According to the quoted vLLM announcement, Inkling-Small features 276B total parameters with 12B active. It natively supports text, image, and audio inputs, and boasts a 1M-token context window. The open-weight model also ships with Day 0 vLLM support optimized for the Blackwell architecture.
Related event: Thinking Machines Releases Open-Source Inkling-Small Model(19 posts)→
More from Infra
- Open Source Project Logs Hidden LLM Serving Traps — alexcovo_eth · 2026-07-31
- Would 10k tok/s Decode Speed Unlock New LLM Use Cases? — LivingSwitch · 2026-07-31
- Inference Optimizations Yield 10x Gains, GPUs May Echo Dark Fiber Lesson — chandan1_ · 2026-07-31
- Stanford Researchers Propose Measuring AI Efficiency by 'Intelligence Per Watt' — StanfordAILab · 2026-07-31
- New Podcast Episode Asks: Do Data Centers Consume a Lot of Water? — AndyMasley · 2026-07-31
- DeepSeek Plans 1GW AI Data Center in Inner Mongolia, Anthropic Targets 9-10GW by 2027 — zephyr_z9 · 2026-07-31