MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model
TheKanter · x · 2026-09-25
MLCommons announced that MLPerf Training v6.1 (October 2026 round) introduces the suite's first LLM post-training benchmark: measuring how fast systems can train a 397B-parameter open-weight model to repair real software, scored on pass@4 quality rather than throughput alone.
Key details:
- Rationale: post-training has driven most LLM gains since H2 2025 (e.g., Qwen 3.8 27B nearing frontier, GLM-5.3's generational leap), yet MLPerf previously only covered pre-training.
- Workload: RLVR — a coding agent solves agentic SWE tasks in a sandbox; multiple rollouts per problem get binary pass/fail rewards, then GRPO updates the policy within each group.
- Target: a full real-world post-training workload under 1024 GB300 hours.
It's the first public, comparable yardstick for post-training efficiency across hardware stacks.
More from Infra
- Former Intel CEO calls HBM "lousy" at Hot Chips 2026 as High Bandwidth Flash looms — Glittering_Depth_722 · 2026-09-25
- kvcached brings virtual memory to LLM KV cache, deployed on 10K+ GPUs — techNmak · 2026-09-25
- IEEE plenary talk: micro-optimizations across the full stack, from silicon to models — fooobar · 2026-09-25
- Dev builds local AI GTM workflow, argues the next platform entry point is hardware-bound — dotey · 2026-09-25
- US Faces Memory Chip Conundrum as AI-Critical Prices Skyrocket, WSJ Reports — pstAsiatech · 2026-09-25
- Qwen Flash Next IQ4_XS beats 27B FP8 on MMLU-Pro, GPQA and GSM8K in community eval — smallDeltaBigEffect · 2026-09-25