K2 Horizon's frozen-model LoRA gives 3x faster inference, trained on 20T tokens
rohanpaul_ai · x · 2026-09-11
K2 Horizon ships with Uno Diffusion, a LoRA adapter that leaves the original autoregressive model frozen while learning to generate blocks of tokens in parallel — no separate draft model, no base-model migration, and IFM reports roughly 3x faster inference with no quality loss.
Other details:
- Each model pretrained on 20T tokens mixing web, code, math, science, multilingual and synthetic data
- 17% of the pre-training corpus contains explicit reasoning trajectories; 10T synthetic tokens used
- Post-training involved 100M+ unique generated tasks
- The whole fleet shares one training, chat, tool-calling and deployment stack, with day-zero vLLM, SGLang and Ollama support across NVIDIA, AMD and Cerebras hardware
Related event: K2 Horizon trained on 20T tokens, LoRA parallel decoding triples speed(2 posts)→
More from Infra
- Ollama adds ChatGPT integration, letting users mix local and cloud models in one app — ollama · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- vLLM upgrade guide: KV offloading, queue admission control, 33.6% Blackwell latency cut — vllm_project · 2026-09-11
- vLLM v0.29.0 cuts Blackwell E2E latency 33.6%, with 6.6-7.6x kernel speedups for Kimi-K3 — vllm_project · 2026-09-11
- One cheeseburger emits as much CO2 as 63,000 Gemini text prompts, math shows — recallingmemories · 2026-09-11