Photon 2.0 targets physical AI inference and claims 2.3x throughput over vLLM
MikeBirdTech · x · 2026-08-04
Photon 2.0 targets physical AI inference workloads
Moondream says Photon 2.0 is a new inference engine built for physical AI, supporting Moondream, Qwen, and Gemma today, with more models coming.
What it claims
- Photon 2.0 is designed for workloads like robots, cameras, industrial systems, and computer-use agents.
- The team says these workloads differ from chat: they have low concurrency, stricter latency needs, and often serve live perception.
- In matched benchmarks on NVIDIA H100, Photon reportedly beats vLLM and SGLang on every throughput test.
- The company says it can reach up to 2.3× higher throughput and bring models online faster.
Why the company says a new engine is needed
- Existing inference stacks are optimized for batching prompts and maximizing tokens per second across GPU fleets.
- Physical AI, by contrast, continuously ingests images, audio, and video and often sacrifices throughput for response time.
- Moondream argues the combination of model × chip × deployment objective is too varied for hand-tuned kernels to scale well.
Bigger picture
Photon 2.0 is positioned as part of the infrastructure stack for deploying perception-heavy AI systems across edge, on-prem, and cloud environments.
Related event: Moondream Launches Photon Inference Compiler, Throughput Up to 2.33x Faster(6 posts)→
More from Infra
- Exa says its web index has 80B pages and is on track for Google-scale in 2027 — garrytan · 2026-08-04
- OpenCode Go says it processed 6T tokens in a single day, led by DeepSeek models — ycombinator · 2026-08-04
- Cloud and AI vendors are creating expensive lock-in and uncontrolled future risks — DavidLinthicum · 2026-08-04
- Swarms Cloud adds saved workflows, session persistence, and 1,500+ models — KyeGomezB · 2026-08-04
- SK Hynix prepares to break ground on $3.87B Indiana HBM packaging fab — rwang07 · 2026-08-04
- Kimi K3 reportedly runs on an 8GB CPU setup by streaming experts from SSD — porAssass · 2026-08-04