Self-Hosting Qwen3.8-27B NVFP4 on SGLang Hits 200+ Tokens/Sec

gnukeith · x · 2026-08-15

Developer @gnukeith keeps expanding his self-hosted stack: Vaultwarden, SGLang, Qwen3.8-27B NVFP4, Open WebUI, Forgejo, FUTO Notes Sync, and Immich.

The highlight: SGLang serving Qwen3.8-27B NVFP4 runs at roughly 200+ tokens per second, making local inference feel smooth.

Original post →

More from Infra

Infra channel →