K3 2.8T Model Hits 947 tokens/s Decoding on Single B300 Node
casper_hansen_ · x · 2026-08-03
Casper Hansen shared decoding speed benchmarks for the 2.8T parameter K3 model running on a single B300 node.
- Overall Throughput: Achieved 947 tokens/s decoding throughput at batch size 32.
- Single User Experience: Delivered 152 tokens/s for a single user.
This is considered highly efficient for a 2.8 trillion parameter model operating on a single node.
More from Infra
- Minimax H3 Tested: Runs Locally on 8GB VRAM — inuptia · 2026-08-04
- Compute Scarcity vs. Creativity: Debating the Future of Neo AI Labs — reneeshah123 · 2026-08-04
- Self-Improving Agents Optimize vLLM, Boosting DeepSeek Throughput by 16% — yisongyue · 2026-08-04
- NVIDIA and KAIST Launch Joint AI Lab to Advance Agentic AI in Korea — hyunw_kim · 2026-08-04
- Running Frontier Models on 24GB VRAM: Local Deployment Challenges Cloud — mintybadgerme · 2026-08-04
- Self-Hosting AI Dev Environments: Sandboxing and Multi-Model Orchestration — Illhoon · 2026-08-04