Temp 0 doesn't guarantee deterministic LLM output — batch shape and kernels shift logits

JFPuget · x · 2026-10-07

In a discussion started by JFPuget, AIQuanting explains why temperature=0 doesn't make LLM inference deterministic: it only removes sampling, while batch shape, kernel reduction order, and which GPU the job lands on all still move the logits. Floating-point parallelism makes bit-exact reproducibility impossible from temp alone.

Original post →

More from Infra

Infra channel →