Why temp=0 LLM inference still isn't deterministic: floating-point order and parallelism

ducha_aiki · x · 2026-10-08

Explaining why temperature=0 doesn't make LLM inference deterministic: hardware/software optimize for speed, not determinism. Floating-point results depend on summation order—(a+b)+c can differ from a+(b+c) due to finite precision, and parallelism makes order unguaranteable. On top of sampling removal, batch shape, kernel reduction order, and which GPU you land on all shift the logits.

Related event: Why LLM Inference Stays Nondeterministic Even at temperature=0(3 posts)→

Original post →

More from Infra

Infra channel →