Cognition Sees 4.8x Token Throughput on Vera Rubin NVL72 as CoreWeave Brings It to Production
NVIDIA Blog · rss · 2026-09-30
At its Fully Connected event, CoreWeave announced availability of NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet, with Cognition (the lab behind Devin) as the first customer running production workloads.
- Performance: benchmarked on a real-world workload sampled from FrontierCode, Vera Rubin NVL72 delivered up to 4.8x total token throughput vs GB200 NVL72 for SWE-2 inference.
- NVIDIA Vera CPU: purpose-built for agents with 128 CPUs / 11,264 cores per rack supporting 11,000+ concurrent isolated environments; CoreWeave measured 3x faster agent sandbox startups and a 1.7x gain on Terminal-Bench.
- CoreWeave Forge: a connected training-eval-improvement environment integrating Weights & Biases, OpenPipe post-training, and open-source marimo; includes ARIA (GA), Agent Lens (20% better failure detection), GA Sandboxes, and serverless RL that's 1.4x faster at 40% lower cost. Canva, Capital One, and MasterClass are early adopters.
- NVIDIA's Ian Buck noted CoreWeave's V100s are still serving customers nearly a decade after Volta; CoreWeave also announced reserved RTX PRO 6000 capacity for healthcare provider Ennoble Care.
More from Infra
- Delip Rao: Most big-budget GPU training runs are run sub-optimally — deliprao · 2026-10-01
- Tencent Leases 100,000 Chips From Oracle, Ex-OpenAI Exec Calls It Insane — Miles_Brundage · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01
- WUSH-KV: Data-Adaptive Transforms for 2-bit KV-Cache Quantization Integrated into SGLang — ISTA-DASLab · 2026-10-01
- Linewise's video agent hits 31x GPU throughput at 1/15 cost on Inco inference infra — songhan_mit · 2026-10-01
- Memory stocks rally overnight in Asia as AI demand frenzy reignites — firstadopter · 2026-10-01