K2 Horizon's frozen-model LoRA gives 3x faster inference, trained on 20T tokens

rohanpaul_ai · x · 2026-09-11

K2 Horizon ships with Uno Diffusion, a LoRA adapter that leaves the original autoregressive model frozen while learning to generate blocks of tokens in parallel — no separate draft model, no base-model migration, and IFM reports roughly 3x faster inference with no quality loss.

Other details:

Related event: K2 Horizon trained on 20T tokens, LoRA parallel decoding triples speed(2 posts)→

Original post →

More from Infra

Infra channel →