Temperature 0 doesn't make LLMs deterministic: 1,000 runs of Qwen3-235B yielded 80 outputs
lmoroney · x · 2026-10-08
Laurence Moroney answers a common question: does temperature=0 make an LLM give the same answer every time? In practice, usually not.
- Temperature 0 means greedy decoding, but most output drift comes from the serving side
- Floating-point addition results depend on summation order, and many GPU kernels change how they add things when batch size changes — which in turn depends on how busy the server is
- Thinking Machines Lab measured this in 2025: 1,000 temperature-0 completions of the same prompt on Qwen3-235B produced 80 different outputs (identical for the first 102 tokens); switching vLLM to batch-invariant kernels made all 1,000 identical
- Takeaway: apps testing LLM features should tolerate small output drift
Related event: Why temperature=0 still can't make LLM inference deterministic(4 posts)→
More from Research
- MarODE, a Markovian ODE framework for scoring LLM reasoning traces, accepted at TMLR — Tanmoy_Chak · 2026-10-08
- Empirical eval shows AI-written unit and integration tests don't improve agent success rates — GabGarrett · 2026-10-08
- OpenAI researcher speculates math AI could soon prove P != NP for good — markjeffrey · 2026-10-08
- Cryptographer: AI doing amazing math puts elliptic curve hardness in no danger — markjeffrey · 2026-10-08
- macro2mind Trains LLMs on Prediction Markets to Simulate Individuals, +15.5 Points Zero-Shot — youjiaxuan · 2026-10-08
- FractAL introduces soft acquisition-strategy selection for batch-mode active learning — anshulkundaje · 2026-10-08