LLM-42 Paper at SOSP 2026 Brings Deterministic LLM Inference via Verified Speculation
tianyin_xu · x · 2026-09-30
A new paper, LLM-42: Enabling Determinism in LLM Inference with Verified Speculation (Raja Gond, Aditya K Kamath, Ramachandran Ramjee, Ashish Panwar), will appear at SOSP 2026 in Prague; paper and code are public.
- Problem: same prompt can yield different outputs across runs, caused by floating-point non-associativity plus dynamic batching and GPU kernels whose reduction orders vary with batch size. Disabling dynamic batching kills throughput; batch-invariant kernels couple determinism to kernel design with fixed overhead.
- Approach: inspired by speculative decoding, LLM-42 decodes via a non-deterministic fast path and enforces determinism with a lightweight verify-rollback loop — the verifier replays candidate tokens under a fixed-shape reduction schedule, committing consistent tokens and rolling back the rest.
- Payoff: mostly reuses existing kernels and only pays for workloads that actually need determinism.
More from Infra
- Ben Lorica: AI's Data Problem Moved Downstream — Usability, Not Scarcity, Is the Bottleneck — bigdata · 2026-09-30
- WhiteMatter: All-to-All Cross-Layer KV Sharing Matches Bigger Models With Half the Cache — INK-USC · 2026-09-30
- 24GB barely fits 27B: local LLM fans note VRAM gap on consumer cards — draginol · 2026-09-30
- AMD boosts Radeon iGPU AI/LLM performance up to 23% with Linux 7.4 drivers — Fcking_Chuck · 2026-09-30
- HBM to hit $100B by 2027, memory supply crunch may last through 2028 — Beth_Kindig · 2026-09-30
- Public data center CEO: compute shortage to last years, warns of GPU blackouts — BenBajarin · 2026-09-30