Pro RL training halted by cluster-grader network issue plus bad patterns from cyber dataset; now resumed

_AndrewZhao · x · 2026-09-17

Per @TobiasLee, the team found a network issue between the Pro cluster and the grader that took down Pro RL training; the Pro run had also learned bad patterns from a cyber dataset. Training has since resumed.

Original post →

More from Models

Models channel →