you.com search enters RL loop to train Nemotron, lifting BrowseComp to 45.45%
PolarBearby · x · 2026-10-07
you.com partnered with NVIDIA and CoreWeave to integrate web search into the reinforcement learning loop for the open-weight model Nemotron 3.5 Lightning, post-training it for web search.
Key results:
- BrowseComp accuracy jumped from 36.97% to 45.45%
- Tool calls dropped by 30%, meaning a smarter model that searches less
Under the hood, CoreWeave's RL Rollouts hot-load new weights into live deployments without touching in-flight requests — about 15x faster than a redeploy cycle — making inference-in-the-training-loop RL post-training practical. It's a full-stack demonstration of training open-weight models with external search tools via RL.
More from Infra
- AMD ships ROCm 10.1, targeting storage-to-GPU data movement as the new training bottleneck — AccBalanced · 2026-10-07
- Crusoe's Path: From Stranded-Gas Bitcoin Mining to Prefab Gigawatt-Scale GPU Datacenters — AccBalanced · 2026-10-07
- Intel to keep working with Elon Musk on Terafab push into cutting-edge chips — SumitGup · 2026-10-07
- Musk bets every 5GW of added US power equals roughly 1% GDP growth — XFreeze · 2026-10-07
- Burn 0.22 released: biggest Rust DL framework update yet, 6-15x faster rebuilds — JosephJacks_ · 2026-10-07
- Async training credit assignment called a key 'foomy' breakthrough for RSI — teortaxesTex · 2026-10-07