New Model Hits Opus-Level Benchmarks at Wild Efficiency, RL Infra Details Emerge
nrehiew_ · x · 2026-09-11
nrehiew calls a new model's benchmark numbers "insane" — Sol/Opus-level performance at that efficiency. His thread digs into its RL training infra: dispatch strategies eliminating long-tail stalls, router replay from previous checkpoints, dataset-level caps plus discard schemes to handle short completions, off-policy ratio bounding with loss masking, persistent KV/router caches, and full-vocab OPD on 40+ teacher models at the final stage.
More from Infra
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11
- Eric Schmidt: AI may hit a money wall before a power wall — $1T capital needed — rohanpaul_ai · 2026-09-11
- SpaceX signs another AI compute deal: $1.11B per month, on track for $100B ARR — NinaDSchick · 2026-09-11
- Carmack: Jetson Thor's 128GB at 273GB/s is over-provisioned for real-time robotics — ID_AA_Carmack · 2026-09-11
- YC Demo Day startup touts ultra-pure diamond wafers for data centers, $160M in LOIs — ycombinator · 2026-09-11