NVIDIA Details GPT-6 Astra Ultrafast: Up to 8x Faster Tokens on Blackwell
BenBajarin · x · 2026-10-02
NVIDIA published a blog post explaining how its Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast, now available via the OpenAI API and to eligible ChatGPT Work and Codex users.
- Performance: Through continuous inference optimizations tapping Blackwell's capabilities, Ultrafast generates tokens up to 8x faster than Astra Standard mode.
- Where it matters: coding agents' edit-test-debug loops, reduced latency between tool calls, and more responsive interactive applications — the time-sensitive cycles where agents write code, call tools, check results, and decide next steps.
- Behind the scenes: OpenAI inference lead Philippe Tillet said NVIDIA's tooling and documentation investment lets models write high-performance kernels for Blackwell and Rubin GPUs, optimizing across the full frontier of latency, throughput, and cost.
- OpenAI keeps improving post-deployment performance, using its own models to refine inference.
Related event: NVIDIA Details How Blackwell GPUs Power GPT-6 Astra Ultrafast's 8x Speedup(3 posts)→
More from Infra
- MiniMax H3 with 360 orbit LoRA runs on 8GB VRAM: 736x576 in 7 minutes — big-boss_97 · 2026-10-02
- NVIDIA applied DL VP: at the scaling limit, efficiency is the new intelligence — ctnzr · 2026-10-02
- IBM's Torch Spyre team: an agentic CI pipeline to keep pace with PyTorch — lmoroney · 2026-10-02
- ModelHarbor: An Open-Source Self-Hosted Model Library — toks-love · 2026-10-02
- US egocentric cleaning data (100s of hrs/week capacity) can't even sell at break-even pricing — paigeinsf · 2026-10-02
- Buyers now fill out export control declarations when purchasing RTX 5090s in stores — blelbach · 2026-10-02