Perplexity's three-layer serving stack: Ivy HTTP, Tulip gRPC, ROSE engine

perplexity_ai · x · 2026-09-05

Perplexity summarizes its standardized inference API stack in three layers: Ivy (Rust HTTP) handles parsing, tokenization, and templating on the CPU and translates to gRPC; Tulip (Rust gRPC) schedules batches for the ROSE engine; ROSE (Python) implements forward passes and CUDA graph management.

Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→

Original post →

More from Infra

Infra channel →