Shibaura Tech's Ozaki Scheme II lands in CUDA 13.4, squeezing FP64 from AI-focused GPUs
udmrzn · x · 2026-10-01
Ozaki Scheme II, a numerical method by Prof. Katsuhisa Ozaki at Shibaura Institute of Technology, has been adopted into NVIDIA's CUDA Toolkit 13.4 (Sept 2026), following the original scheme's 2025 inclusion.
Background: GPU double-precision (FP64) performance has collapsed as chips optimize for AI low-precision math — B200 offers 40 TFLOPS FP64 while Blackwell Ultra (B300) drops to 1.2–1.3 TFLOPS, a 30x decline.
How it works: The scheme mathematically combines multiple low-precision operations (error-free transformations, modular arithmetic via the Chinese Remainder Theorem) to reproduce double-precision accuracy in software, running on Ampere-and-later GPUs with no extra hardware — repurposing idle low-precision units for scientific computing.
Impact: Useful for weather forecasting, seismic design, quantum chemistry and materials science. RIKEN lists it among methods for the Fugaku NEXT supercomputer, and it's reportedly used in US DOE science programs.
More from Infra
- Nvidia authorizes $150 billion buyback, a vote on future AI demand — YvesMulkers · 2026-10-01
- GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support — ycombinator · 2026-10-01
- Pareto partners with Engy to bring Bittensor SN10 inference optimization to external customers — markjeffrey · 2026-10-01
- Respan launches Span-01 router: 37% cheaper than best single model at matching accuracy — ycombinator · 2026-10-01
- RTX 3090 power can be dialed down to 120W for inference, saving power and heat — QuixiAI · 2026-10-01
- AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton — Azaliamirh · 2026-10-01