Inkling Gets Day-One FP4 Inference Support
zhyncs42 · x · 2026-07-16
LightSeek's TokenSpeed has provided day-one FP4 inference support for Thinking Machines Lab's newly released Inkling, covering NVIDIA GB200/GB300's NVFP4 and AMD MI350X/MI355X's MXFP4.
The post emphasizes that this is a unified inference kernel solution built on PyTorch infrastructure. In collaboration with vLLM, the goal is to push high-performance, reusable inference kernels into the open-source AI ecosystem. Implementation highlights include native FP4 serving, a unified kernel architecture, MTP, optimized attention, and heterogeneous KV cache.
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22