Moondream pitches an inference compiler that serves 1.01×–2.33× vLLM and SGLang throughput

suchenzang · x · 2026-08-04

A Moondream post argues for an inference compiler that turns a model description into an optimized megakernel, letting the GPU run the whole inference path while the CPU stays free for other work.

The accompanying benchmark claims Photon matches or beats on ChartQA throughput, with gains across batch sizes up to 8. The screenshot says Photon serves 1.01×–2.33× the throughput of vLLM and SGLang, and the author says faster boot times matter when a robot crashes or reboots and needs its “brain” back immediately.

Related event: Moondream Launches Photon Inference Compiler, Throughput Up to 2.33x Faster(6 posts)→

Original post →

More from Infra

Infra channel →