Inkling Gets Day-One FP4 Inference Support

zhyncs42 · x · 2026-07-16

LightSeek's TokenSpeed has provided day-one FP4 inference support for Thinking Machines Lab's newly released Inkling, covering NVIDIA GB200/GB300's NVFP4 and AMD MI350X/MI355X's MXFP4.

The post emphasizes that this is a unified inference kernel solution built on PyTorch infrastructure. In collaboration with vLLM, the goal is to push high-performance, reusable inference kernels into the open-source AI ecosystem. Implementation highlights include native FP4 serving, a unified kernel architecture, MTP, optimized attention, and heterogeneous KV cache.

Original post →

More from Infra

Infra channel →