Perceptron's Multilook API Prefills Video Context Once, Cuts Input Cost to 32% at 16 Prompts

AkshatS07 · x · 2026-09-05

Multimodal API startup Perceptron launched Multilook, an endpoint that prefills a shared media context once and reuses it across up to 16 prompts per call. At 16 prompts, input costs drop to 32% of sending requests separately; throughput improves 1.6–3.9× under load, and median latency drops 2–4.9×. The trick: a Files API saves uploads but not the repeated forward pass over the same frames—Multilook eliminates the redundant prefills entirely.

Original post →

More from Infra

Infra channel →