NVIDIA CCCL 3.1 adds three-tier FP determinism control; cross-GPU reproducibility costs 20-30%
blelbach · x · 2026-10-07
An NVIDIA technical blog introduces floating-point determinism controls in CUDA Core Compute Libraries (CCCL) 3.1, which developer David Aronchick shared to push back on the claim that GPU parallelism is inherently non-reproducible.
Key points:
- Floating-point addition isn't associative, so parallel reductions vary between runs due to differing summation order — the root cause of non-determinism.
- CCCL 3.1 adds a single-phase CUB API that configures reduction determinism via an execution environment, with three levels:
- notguaranteed: maximum performance, results may vary
- runtorun (default): consistent results on the same GPU
- gputogpu: bitwise-identical results across GPUs using a Reproducible Floating-point Accumulator with three exponent bins, giving tighter error bounds than standard pairwise summation
- Cost: gputogpu adds 20%-30% execution time for large problems on an H200.
- Currently limited to reductions; expansion to more CUDA parallel primitives is tracked on GitHub.
Aronchick argues non-determinism in some implementations is a gap, not a fundamental design flaw, and can now be controlled explicitly.
Related event: Debating GPU Non-Determinism as NVIDIA CCCL 3.1 Adds Deterministic Controls(3 posts)→
More from Infra
- Team claims sub-5-second full weight sync for 1T-parameter RL training — saurabh_shah2 · 2026-10-07
- NVIDIA's UNREAL: one model unifies corpus retrieval and long-context at 128K+ — nvidia · 2026-10-07
- SlimWise prunes MoE experts only at decode, boosting throughput up to 1.81x — Gunho Park · 2026-10-07
- Ora scanned 107,797 sites: average agent-readiness score just 41/100 — EdenEmarco177 · 2026-10-07
- Solo dev hits Together AI rate limits running parallel agents, seeks generous API credits — Correct_Positive_108 · 2026-10-07
- Ramjet: an open-source local alternative to NVIDIA Dynamo for multi-GPU inference — DoggoProfessor959 · 2026-10-07