Ornith-1.5 DFlash draft models deliver up to 2.54x lossless inference speedup

alan_ritter · x · 2026-10-01

Ornith-1.5 (9B, 35B-A3B, 397B) now ships with DFlash draft models, available on Hugging Face under the Ornith-1.5-DFlash collection (draft sizes 1B, 0.4B, 1B).

The approach is speculative decoding: DFlash drafts multiple tokens at once, then Ornith-1.5 verifies them in a single pass instead of generating one token at a time. Official benchmarks show up to 2.54x speedup with no loss in generation quality. Users can plug the draft models into their local workstations.

Related event: Ornith-1.5 Open-Sources DFlash Draft Models for Up to 2.54x Lossless Speedup(2 posts)→

Original post →

More from Infra

Infra channel →