Thinky publishes model card with B200 benchmark results and SGLang vLLM support

TheZachMueller · x · 2026-07-22

Thinky Machines has published its model card, after waiting for SGLang and vLLM support to land so the numbers would be meaningful.

The benchmark in the image was run on an NVIDIA HGX B200 with CUDA 13.0, using the official NVFP4 checkpoint at TP=8. The model is described as fast and capable, and the card reports throughput figures including 2,510.64 tok/s generation throughput per user and 12,553.21 tok/s total throughput, with 2,023.21 ms mean TTFT and 11.76 ms mean ITL.

Related event: Thinky Releases Model Card with B200 Benchmark Data(2 posts)→

Original post →

More from Models

Models channel →