Photon 2.0 compiles Moondream, Qwen 3.5 and Gemma 4 into megakernels
sloppenheimer · x · 2026-08-04
Photon 2.0 is presented as a compiler-driven inference engine that turns Moondream, Qwen 3.5, and Gemma 4 into “megakernels” — a single GPU program for the full forward pass.
The launch claims throughput gains over vLLM and SGLang, with the chart in the post showing improvements ranging from +10.4% to +132.7% on NVIDIA H100 under the cited benchmark setup. The pitch is that physical-AI workloads need a dedicated inference stack rather than hand-tuned kernels.
More from Infra
- Wan 2.1 now runs locally on supported Samsung phones via Saient Quartz — SaientAI · 2026-08-04
- NVIDIA pitches agentic commerce for retail, with merchant-controlled checkout and pricing — nvidia · 2026-08-04
- Stripe Projects lets AI agents add hosting, auth, databases, and billing from the CLI — jeff_weinstein · 2026-08-04
- Multi-agent workflows can burn billions of tokens unless you control duplication — HaktanSuren · 2026-08-04
- Next.js 16.3 cuts dev RAM by 90% and adds docs for coding agents — cramforce · 2026-08-04
- TokTier speeds up agent serving with exact stateful tokenization and stable-boundary repair — omarsar0 · 2026-08-04