Perplexity's ROSE engine: ragged attention, no KV cache for embeddings

perplexity_ai · x · 2026-09-05

Perplexity details ROSE, its model engine: it reuses the same kernels for LLMs and embeddings, but for embeddings it skips the KV cache and uses ragged attention instead of paged attention. ROSE supports multiple attention backends, with kernel choice depending on model shape and sequence length.

Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→

Original post →

More from Infra

Infra channel →