New Model Hits Opus-Level Benchmarks at Wild Efficiency, RL Infra Details Emerge

nrehiew_ · x · 2026-09-11

nrehiew calls a new model's benchmark numbers "insane" — Sol/Opus-level performance at that efficiency. His thread digs into its RL training infra: dispatch strategies eliminating long-tail stalls, router replay from previous checkpoints, dataset-level caps plus discard schemes to handle short completions, off-policy ratio bounding with loss masking, persistent KV/router caches, and full-vocab OPD on 40+ teacher models at the final stage.

Related event: DeepSeek V4.1 Tech Report Deep Dive: RL Infrastructure, Sandbox Design and Inference Stack(8 posts)→

Original post →

More from Infra

Infra channel →