Free Qwen3.8-Flash-Next endpoint launched on 4× H200 at 100+ tok/s

victormustar · x · 2026-08-26

Developer victormustar deployed a free public endpoint for Qwen3.8-Flash-Next, requiring no API tokens. It is OpenAI-compatible, supports vision, tool calls, and 262K context. Powered by 4× H200 GPUs (FP8) via SGLang, it achieves 140 tok/s per stream (100 tok/s @ 16 concurrent) with 0.8s TTFT.

Original post →

More from Infra

Infra channel →