Photon 2.6 ships FP8 + speculative decoding, runs Qwen3.5 27B at 400+ tok/s on B200

Bedrovelsen · x · 2026-09-30

Photon 2.6 is out with FP8 inference and speculative decoding support. The author demonstrates running Qwen3.5 27B at over 400 tokens per second on an NVIDIA B200, showcasing the throughput of the latest inference stack on flagship accelerators.

Original post →

More from Infra

Infra channel →