MLPerf v6.1 debuts agentic benchmark and first 1T+ parameter model run
TheZachMueller · x · 2026-09-16
Lambda's MLPerf Inference v6.1 results mark two firsts: the first agentic inference workload on datacenter hardware and the first deployment of a 1T+ parameter model (Kimi K2.6 on HGX B200, replacing the 27B reference). On 4× Blackwell Ultra, Lambda posted the leading Offline throughput for GPT-OSS 120B (65,511 tokens/s) plus an 8.85% throughput gain over v6.0 on identical hardware from pure software optimization.
More from Infra
- Even Preemptible 8xH100 Instances Are Sold Out Everywhere — bingxu_ · 2026-09-16
- Custom Midtrain Plus Own RL Matches Astra Max at Half the Inference Price — hsu_byron · 2026-09-16
- Apple A20 Pro die shot revealed: 8.00x12.35mm die exposed in chip photo library — BenBajarin · 2026-09-16
- AI Infra Summit: scaling AI compute poses far deeper engineering challenges than most realize — BenBajarin · 2026-09-16
- Argentina's La Nacion covers a millions-strong miscount, spotlighting GeoParquet and DuckDB — MaxLenormand · 2026-09-16
- 27B uncensored Qwen at 160K context on one RTX 5090: 140-190 tok/s with DFlash2 — Fz1zz · 2026-09-16