Thinking Machines explains why LLM inference stays nondeterministic even at temperature 0
JFPuget · x · 2026-10-07
Horace He of Thinking Machines Lab published "Defeating Nondeterminism in LLM Inference," digging into why LLM outputs are hard to reproduce.
The problem: Even with temperature set to 0 (greedy sampling, theoretically deterministic), outputs still vary — both on LLM APIs like ChatGPT and on self-hosted inference with vLLM or SGLang.
Questioning the popular hypothesis: The widely cited "concurrency + floating point" explanation — GPU floating-point non-associativity plus racing concurrent cores changing execution order — is examined in depth, with the post exploring what actually breaks determinism and how to fix it.
A must-read for developers and researchers who depend on reproducible LLM outputs.
Related event: Why temperature=0 still can't guarantee deterministic LLM inference(2 posts)→
More from Infra
- Rumor: chip startup Terafab in talks with all three logic foundries, Samsung ahead — pstAsiatech · 2026-10-07
- Adaption AI launches AutoScientist Leaderboard ranking custom models across 44 domains — sarahookr · 2026-10-07
- CoreWeave enters India with 240 MW AdaniConneX deployment in Navi Mumbai — ayushthakur0 · 2026-10-07
- Qdrant compresses Google EmbeddingGemma 2 vectors 77x with only 5% quality loss — qdrant_engine · 2026-10-07
- ASML tipped to ship 120-125 EUV tools in 2028 as capacity expansion accelerates — zephyr_z9 · 2026-10-07
- Local 27B vs Cloud API: 89.5 vs 92.6 Quality, Fully Green Coding Runs — Icy-Stay-1004 · 2026-10-07