Ornith-1.5 DFlash draft models deliver up to 2.54x lossless inference speedup
alan_ritter · x · 2026-10-01
Ornith-1.5 (9B, 35B-A3B, 397B) now ships with DFlash draft models, available on Hugging Face under the Ornith-1.5-DFlash collection (draft sizes 1B, 0.4B, 1B).
The approach is speculative decoding: DFlash drafts multiple tokens at once, then Ornith-1.5 verifies them in a single pass instead of generating one token at a time. Official benchmarks show up to 2.54x speedup with no loss in generation quality. Users can plug the draft models into their local workstations.
More from Infra
- Paraguay's $12B Iguazú AI City to host Latin America's largest data center — teortaxesTex · 2026-10-01
- You.com, NVIDIA and CoreWeave bring live web search into RL training, starting with Nemotron 3.5 Lightning — RichardSocher · 2026-10-01
- Padding trick lets vLLM run tp=6: 27B model at 50 tok/s on six 7900 XTX GPUs — Biomass23 · 2026-10-01
- Chutes team on Bittensor Subnet 64 may have stumbled on a new way to train models while fixing inference economics — markjeffrey · 2026-10-01
- Manager locked Teams transcripts, employee used Copilot to dig JSON URL out of page source — TheBigCrowbroski · 2026-10-01
- Compute per MW comparison: Nvidia still best price/perf despite prices — Storge2 · 2026-10-01