Free Qwen3.8-Flash-Next endpoint launched on 4× H200 at 100+ tok/s
victormustar · x · 2026-08-26
Developer victormustar deployed a free public endpoint for Qwen3.8-Flash-Next, requiring no API tokens. It is OpenAI-compatible, supports vision, tool calls, and 262K context. Powered by 4× H200 GPUs (FP8) via SGLang, it achieves 140 tok/s per stream (100 tok/s @ 16 concurrent) with 0.8s TTFT.
More from Infra
- Local AI Registry: Open source index for hardware, models, and deployment recipes — StefanoGogioso · 2026-08-26
- TRANSIT runtime cuts LLM training GPU needs by up to 50% — PyTorch · 2026-08-26
- Ollama Announces GLM-5.3-Flash Coming Soon to Cloud Service — ollama · 2026-08-26
- Polymarket: 13% Chance AI Bubble Bursts by End of 2026 — Polymarket · 2026-08-26
- Open-Source GPU Price Aggregator Vram Watch Released — KyeGomezB · 2026-08-26
- OpenAI plans world's largest data center in Ohio, requiring more power than all state homes — bennash · 2026-08-26