Perplexity's Tulip: CUDA graphs and LazyTensors for low-latency batch embedding
perplexity_ai · x · 2026-09-05
Perplexity explains Tulip, its lightweight Rust gRPC inference server between Ivy and the ROSE engine: it batches incoming requests for the GPU, noting that for small embedding models runtime depends on tokens (512 tokens fills the GPU). Two techniques cut latency and boost throughput: CUDA graphs pre-record GPU work for single-call launches, and LazyTensors tracks results asynchronously so the CPU preps the next batch while the GPU finishes the current one.
Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→
More from Infra
- Burn Bar for Omarchy visualizes Claude/Codex token burn, quotas and GPU load locally — DanWahlin · 2026-09-05
- Scaling wall? Reddit argues test-time compute is the industry's new playbook — erdematar · 2026-09-05
- Tencent Hunyuan Hy4 preview: 770B total/49B active, 1M context, Apache 2.0, day-0 vLLM — aftahi_ai · 2026-09-05
- Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate — ChopSticksPlease · 2026-09-05
- Speechify CEO on self-built data centers, ElevenLabs leapfrog, and the $15M AI talent war — 20VC · 2026-09-05
- He Uses Local LLMs Like a 3D Printer: 12 Games, 29 Mods and Countless Tools Built Solo — Quebber · 2026-09-05