Perceptron's Multilook API Prefills Video Context Once, Cuts Input Cost to 32% at 16 Prompts
AkshatS07 · x · 2026-09-05
Multimodal API startup Perceptron launched Multilook, an endpoint that prefills a shared media context once and reuses it across up to 16 prompts per call. At 16 prompts, input costs drop to 32% of sending requests separately; throughput improves 1.6–3.9× under load, and median latency drops 2–4.9×. The trick: a Files API saves uploads but not the repeated forward pass over the same frames—Multilook eliminates the redundant prefills entirely.
More from Infra
- After ChatGPT, Claude & Grok All Went Dark, One User's 3-Machine Local AI Lab Kept Running — cocktailpeanut · 2026-09-05
- DRAM density has flattened: servers stuck at 8TB for 5 years, CXL is the way out — lauriewired · 2026-09-05
- Benchmarked 21 Qwen3.8-27B quants on 16GB VRAM: bartowski IQ4_XS wins — Storterald · 2026-09-05
- Nvidia DLSS 5 frame interpolation discussed in Stable Diffusion community — KonoTheSavage1 · 2026-09-05
- Dylan Patel on Dwarkesh: How Elon Musk Played the Compute Market — Dwarkesh Patel · 2026-09-05
- MiniMax and Together AI host London event on the economics of open-model production AI — MiniMax_AI · 2026-09-05