Thinky publishes model card with B200 benchmark results and SGLang vLLM support
TheZachMueller · x · 2026-07-22
Thinky Machines has published its model card, after waiting for SGLang and vLLM support to land so the numbers would be meaningful.
The benchmark in the image was run on an NVIDIA HGX B200 with CUDA 13.0, using the official NVFP4 checkpoint at TP=8. The model is described as fast and capable, and the card reports throughput figures including 2,510.64 tok/s generation throughput per user and 12,553.21 tok/s total throughput, with 2,023.21 ms mean TTFT and 11.76 ms mean ITL.
Related event: Thinky Releases Model Card with B200 Benchmark Data(2 posts)→
More from Models
- Gemini 3.6 Flash Becomes the Default Model for Managed Agents — _philschmid · 2026-07-23
- Cohere Confirms Community Quants for Arabic Speech Model on Hugging Face — cohere · 2026-07-23
- Reddit jokes about “obliterating” Kimi K3 the moment it ships — JsonBasedman · 2026-07-23
- Leak: OpenAI Targets September Release for GPT-6 as Autonomous AI Researcher — VraserX · 2026-07-23
- Rumor says GPT-6 has been delayed by several more months — patience_cave · 2026-07-23
- Moonshot’s Kimi K3 is blamed for Friday panic as AI prices keep falling — TiernanRayTech · 2026-07-23