Photon claims up to 2.33× throughput over vLLM and SGLang on H100
AccBalanced · x · 2026-08-04
- Moondream claims Photon delivers 1.01×–2.33× the throughput of vLLM and SGLang in benchmark tests.
- The chart compares request throughput across 1, 2, 4, and 8 concurrent streams on models including Moondream 3, Qwen3.5 variants, and Gemma 4 variants.
- It also highlights a second benefit: faster model boot times, which they say matters for robotics or other systems that need to recover quickly after a crash or reboot.
- The footer lists the tested stack as Photon 2.0.0, vLLM 0.25.1, SGLang 0.5.3, running on NVIDIA H100 80GB HBM3.
Related event: Moondream Launches Photon Inference Compiler, Throughput Up to 2.33x Faster(6 posts)→
More from Infra
- Exa says its web index has 80B pages and is on track for Google-scale in 2027 — garrytan · 2026-08-04
- OpenCode Go says it processed 6T tokens in a single day, led by DeepSeek models — ycombinator · 2026-08-04
- Cloud and AI vendors are creating expensive lock-in and uncontrolled future risks — DavidLinthicum · 2026-08-04
- Swarms Cloud adds saved workflows, session persistence, and 1,500+ models — KyeGomezB · 2026-08-04
- SK Hynix prepares to break ground on $3.87B Indiana HBM packaging fab — rwang07 · 2026-08-04
- Kimi K3 reportedly runs on an 8GB CPU setup by streaming experts from SSD — porAssass · 2026-08-04