Perplexity's Ivy: the HTTP gateway handling tokenization and batch splitting
perplexity_ai · x · 2026-09-05
Perplexity details Ivy, the HTTP gateway in its serving stack: it handles CPU-side request prep (parsing, tokenization, templating) and splits large batches before sending to Tulip over gRPC, letting the team tune request formatting and tokenization without touching the heavier inference servers.
Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→
More from Infra
- Burn Bar for Omarchy visualizes Claude/Codex token burn, quotas and GPU load locally — DanWahlin · 2026-09-05
- Scaling wall? Reddit argues test-time compute is the industry's new playbook — erdematar · 2026-09-05
- Tencent Hunyuan Hy4 preview: 770B total/49B active, 1M context, Apache 2.0, day-0 vLLM — aftahi_ai · 2026-09-05
- Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate — ChopSticksPlease · 2026-09-05
- Speechify CEO on self-built data centers, ElevenLabs leapfrog, and the $15M AI talent war — 20VC · 2026-09-05
- He Uses Local LLMs Like a 3D Printer: 12 Games, 29 Mods and Countless Tools Built Solo — Quebber · 2026-09-05