Perplexity's ROSE engine: ragged attention, no KV cache for embeddings
perplexity_ai · x · 2026-09-05
Perplexity details ROSE, its model engine: it reuses the same kernels for LLMs and embeddings, but for embeddings it skips the KV cache and uses ragged attention instead of paged attention. ROSE supports multiple attention backends, with kernel choice depending on model shape and sequence length.
Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→
More from Infra
- Burn Bar for Omarchy visualizes Claude/Codex token burn, quotas and GPU load locally — DanWahlin · 2026-09-05
- Scaling wall? Reddit argues test-time compute is the industry's new playbook — erdematar · 2026-09-05
- Tencent Hunyuan Hy4 preview: 770B total/49B active, 1M context, Apache 2.0, day-0 vLLM — aftahi_ai · 2026-09-05
- Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate — ChopSticksPlease · 2026-09-05
- Speechify CEO on self-built data centers, ElevenLabs leapfrog, and the $15M AI talent war — 20VC · 2026-09-05
- He Uses Local LLMs Like a 3D Printer: 12 Games, 29 Mods and Countless Tools Built Solo — Quebber · 2026-09-05