OpenAI's GPT-6 Astra Ultrafast Delivers Up to 8x Faster Token Generation on NVIDIA Blackwell
NVIDIA Blog · rss · 2026-10-02
OpenAI launched GPT-6 Astra Ultrafast, accelerated by NVIDIA Blackwell GPUs with inference optimizations delivering up to 8x faster token generation than Astra Standard. It's live in the OpenAI API and for eligible ChatGPT Work and Codex users.
- Why it matters: speed compounds in agent loops — code edit-test-debug cycles, tool-call gaps, and interactive apps all get snappier.
- Ongoing gains: OpenAI uses its internal models to optimize inference software on NVIDIA GPUs, leveraging platform programmability; inference lead Philippe Tillet says Astra turns GPU-programming knowledge into high-performance kernels across latency, throughput, and cost.
- Flexibility: the programmable infrastructure lets teams reuse compute across training, inference, and RL, improving utilization.
Access, pricing, and implementation details are in the Ultrafast guide.
More from Infra
- Redditor builds fully local LLM-powered radio site on two DGX Sparks and a 5090 — jwhh91 · 2026-10-02
- VC quip: many neoclouds are closer to 95% than five nines of reliability — saranormous · 2026-10-02
- Report: lenders demand up to 25% collateral from Nvidia as GPU-backed loans wobble — GaryMarcus · 2026-10-02
- Microsoft Backs Snowflake-Led Effort to Standardize Business Metrics for AI — xiaosun86 · 2026-10-02
- Full SGLang config for GLM-5.3-Flash NVFP4 on 4x RTX 6000 Max-Q — TheZachMueller · 2026-10-02
- GLM-5.3-Flash NVFP4 benchmarks show no per-user speedup beyond 8 concurrent requests — TheZachMueller · 2026-10-02