Photon 2.6 ships FP8 + speculative decoding, runs Qwen3.5 27B at 400+ tok/s on B200
Bedrovelsen · x · 2026-09-30
Photon 2.6 is out with FP8 inference and speculative decoding support. The author demonstrates running Qwen3.5 27B at over 400 tokens per second on an NVIDIA B200, showcasing the throughput of the latest inference stack on flagship accelerators.
More from Infra
- Raja Koduri: Western AI compute costs $50-60B per gigawatt, China targets under $10B — RajaXg · 2026-09-30
- DeepSeek open-sources DeepGEMM Ascend port, hitting 99.8% of hardware limit on GEMM — zheanxu · 2026-09-30
- AI intelligence-cost Pareto frontier shifted fast: GPT-5 mini at 17 ($0.05) to Claude Opus 5.5 at 58 ($5.98) — ArtificialAnlys · 2026-09-30
- Google's Project Suncatcher to fly TPUs in space for the first time on Oct 1 — allisondman · 2026-09-30
- SGLang turns Qwen3.8-27B into a decision model that beats Pokémon FireRed at sub-100ms — zhaoran_wang · 2026-09-30
- Agentic AI turns CPUs into the overlooked bottleneck as CPU:GPU ratios shift upward — AccBalanced · 2026-09-30