Thinking Machines explains why LLM inference stays nondeterministic even at temperature 0

JFPuget · x · 2026-10-07

Horace He of Thinking Machines Lab published "Defeating Nondeterminism in LLM Inference," digging into why LLM outputs are hard to reproduce.

The problem: Even with temperature set to 0 (greedy sampling, theoretically deterministic), outputs still vary — both on LLM APIs like ChatGPT and on self-hosted inference with vLLM or SGLang.

Questioning the popular hypothesis: The widely cited "concurrency + floating point" explanation — GPU floating-point non-associativity plus racing concurrent cores changing execution order — is examined in depth, with the post exploring what actually breaks determinism and how to fix it.

A must-read for developers and researchers who depend on reproducible LLM outputs.

Related event: Why temperature=0 still can't guarantee deterministic LLM inference(2 posts)→

Original post →

More from Infra

Infra channel →