MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model

TheKanter · x · 2026-09-25

MLCommons announced that MLPerf Training v6.1 (October 2026 round) introduces the suite's first LLM post-training benchmark: measuring how fast systems can train a 397B-parameter open-weight model to repair real software, scored on pass@4 quality rather than throughput alone.

Key details:

It's the first public, comparable yardstick for post-training efficiency across hardware stacks.

Original post →

More from Infra

Infra channel →