Orthrus Study: Lossless Speculative Decoding Holds Only at High Numerical Precision
Ilya Koziev · hf · 2026-09-15
A Hugging Face write-up examines how lossless speculative decoding truly is in the Orthrus system. The finding: losslessness only holds under high numerical precision, while BF16 causes trajectory divergence between draft and target outputs. Interestingly, this divergence does not impair downstream benchmark results, suggesting the practical impact of precision-induced drift is limited, though the notion of "lossless" is precision-dependent.
More from Infra
- Subnormal floats are expensive — but only on Intel, benchmarks show — lemire · 2026-09-15
- LLM Inference Engineer dubbed the most AI-proof job by tech commentator — ashishllm · 2026-09-15
- SGLang and Samsung whitepaper: 3.1x lower LLM inference latency via AI Memory Node — ying11231 · 2026-09-15
- Weaviate v1.38 adds Boost: re-rank search results without dropping them — victorialslocum · 2026-09-15
- ACE Step 1.5 Music Generation Now Runs Locally on Android via CPU or OpenCL GPU — sgcego · 2026-09-15
- VMware Private AI Foundation with NVIDIA Tops Private AI Cloud Evaluation at 9.1/10 — DavidLinthicum · 2026-09-15