Temp 0 doesn't guarantee deterministic LLM output — batch shape and kernels shift logits
JFPuget · x · 2026-10-07
In a discussion started by JFPuget, AIQuanting explains why temperature=0 doesn't make LLM inference deterministic: it only removes sampling, while batch shape, kernel reduction order, and which GPU the job lands on all still move the logits. Floating-point parallelism makes bit-exact reproducibility impossible from temp alone.
More from Infra
- POC 2026 talk shows container escape through the NVIDIA GPU driver despite read-only limits — evilsocket · 2026-10-07
- Google Releases SAM, a P2P Network Letting AI Agents Discover and Call Each Other's Tools — thisguyknowsai · 2026-10-07
- Running a 744B MoE on a 25GB laptop: Colibri streams experts from disk like a weight JIT — thisdudelikesAI · 2026-10-07
- Jamie Dimon: Data centers should be built in communities that want them — Famous_Proof_8429 · 2026-10-07
- Virginia Tech's Hybrid Latent Attention boosts looped LLM GPU throughput up to 8.8x with minimal accuracy loss — rohanpaul_ai · 2026-10-07
- Java Vector API: Writing SIMD Directly Since JDK 16 to Unlock Single-Core Performance — lemire · 2026-10-07