Moondream's Photon 2.1 adds streaming ASR, beats vLLM with up to 3.1x throughput
xeophon · x · 2026-09-03
Moondream released Photon 2.1, its realtime multimodal inference engine, adding streaming speech recognition for Whisper, Qwen3-ASR, and Parakeet, plus TTS models and B200 GPU support.
The team reports winning all 16 matched H100/B200 tests against vLLM, Qwen-ASR, and NeMo, with up to 3.1x throughput. Open source, installable via pip install -U moondream.
More from Infra
- MOSTIK tops ARC-AGI leaderboard by piping frontier-model reasoning into small models via latent space — SimplyAnnisa · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03
- FastH3 Now Runs Locally on Apple Silicon and DGX Spark — Vandy_simp · 2026-09-03
- Claim: Without the datacenter buildout campaign, the US would be in a sharp recession — ZeeshanZiaML · 2026-09-03