Astra optimizes its own inference on Rubin chips, doubling throughput in 72 hours
bookwormengr · x · 2026-09-16
At the AI Infra Summit, NVIDIA's Ian Buck and OpenAI's Sachin Katti presented recursive optimization results for Astra on next-gen hardware:
- OpenAI ported a Blackwell-optimized Astra checkpoint to Vera Rubin with minor modifications and got 3x throughput out of the box;
- Astra was then run as an agent to optimize its own inference on Rubin chips;
- Over 72 hours it delivered an additional 2x throughput improvement;
- Work that previously took weeks or months was completed over a weekend.
Attendees called it the kind of capability that 'gives Eliezer Yudkowsky nightmares.'
More from Infra
- Trillion-parameter model trained on own physics lab data hits memory-efficiency SOTA — vwxyzjn · 2026-09-16
- Six efficiency breakthroughs labs didn't see coming upend semiconductor demand assumptions — bookwormengr · 2026-09-16
- Nadella: a 400MW data center grew a rural town's tax revenue 12X — rohanpaul_ai · 2026-09-16
- Developer balks at $16k price for what is essentially a consumer GPU with extra memory — draginol · 2026-09-16
- Cloudflare lets sites block AI training crawlers while staying searchable — djfergus · 2026-09-16
- Hitachi Energy to invest $528 million in new transformer factory in Mississippi — oilmutt · 2026-09-16