Why temp=0 LLM inference still isn't deterministic: floating-point order and parallelism
ducha_aiki · x · 2026-10-08
Explaining why temperature=0 doesn't make LLM inference deterministic: hardware/software optimize for speed, not determinism. Floating-point results depend on summation order—(a+b)+c can differ from a+(b+c) due to finite precision, and parallelism makes order unguaranteable. On top of sampling removal, batch shape, kernel reduction order, and which GPU you land on all shift the logits.
Related event: Why LLM Inference Stays Nondeterministic Even at temperature=0(3 posts)→
More from Infra
- Tracking a 24/7 agent for 30 days: $6 VPS, $22 API, and uptime is the real leak — YamOk7317 · 2026-10-08
- Bain: nearly 150GW of new data center capacity by 2030, requiring $5-6.5 trillion — Beth_Kindig · 2026-10-08
- TypeSafeAI's Jev served a trillion tokens on Modal within three days of launch — josh_wills · 2026-10-08
- llama.cpp PR adds GPU cache for host-resident MoE experts, big speedup potential — jacek2023 · 2026-10-08
- Mapping the firm-power stack behind Google's massive nuclear deal: CEG, TLN, VST, Hubbell, Bloom — demian_ai · 2026-10-08
- NVIDIA releases CUDA 13.4 with new features detailed in official blog — NVIDIAAI · 2026-10-08