Photon 2.0 compiles Moondream, Qwen 3.5 and Gemma 4 into megakernels

sloppenheimer · x · 2026-08-04

Photon 2.0 is presented as a compiler-driven inference engine that turns Moondream, Qwen 3.5, and Gemma 4 into “megakernels” — a single GPU program for the full forward pass.

The launch claims throughput gains over vLLM and SGLang, with the chart in the post showing improvements ranging from +10.4% to +132.7% on NVIDIA H100 under the cited benchmark setup. The pitch is that physical-AI workloads need a dedicated inference stack rather than hand-tuned kernels.

Original post →

More from Infra

Infra channel →