Self-Hosting Qwen3.8-27B NVFP4 on SGLang Hits 200+ Tokens/Sec
gnukeith · x · 2026-08-15
Developer @gnukeith keeps expanding his self-hosted stack: Vaultwarden, SGLang, Qwen3.8-27B NVFP4, Open WebUI, Forgejo, FUTO Notes Sync, and Immich.
The highlight: SGLang serving Qwen3.8-27B NVFP4 runs at roughly 200+ tokens per second, making local inference feel smooth.
More from Infra
- LLMRouter open-sources 16+ implementations with xRouteBench for LLM routing — Justgototheeffinmoon · 2026-08-16
- User Benchmarks Qwen3.8 27b on M5 Max: 8t/s (bf16), 17t/s (8bit) — julianharris · 2026-08-16
- The accountability gap in LLM inference: Proving which model weights actually ran — Some_parts_Bi · 2026-08-16
- WeeLLM: Run FLUX.1-dev on 4GB VRAM Without Quantization — AlarmingPhrase8174 · 2026-08-16
- Enterprise AI buyers push Lenovo to record quarter: revenue up 43%, services margin 3x PC — shashib · 2026-08-16
- Google floats many TPU RFPs, takes more wafers directly to TSMC each generation — BenBajarin · 2026-08-16