Inkling Released with 15% Speed Boost

LysandreJik · x · 2026-07-16

Thinking Machines has launched Inkling: an open large model supporting text, image, and audio inputs with text output, scaled at 1T.

The post highlights the pluggable optimization space for inference acceleration:

The core takeaway is that the kernels approach allows model developers to quickly integrate better implementations just like changing a config, while kernel developers focus on delivering the fastest versions, ultimately benefiting everyone.

Related event: Thinking Machines launches Inkling, its first open-weight multimodal model(169 posts)→

Original post →

More from Infra

Infra channel →