Dev Plans Free Community Access to Local Qwen 3.8 27B at 56 Tokens/sec
No_Run8812 · reddit · 2026-08-18
A developer plans to offer free community access to a locally hosted Qwen 3.8 27B (8-bit quant). Running on a single RTX PRO 6000 96GB via vLLM, it achieves 56 tk/s decode with 3 concurrent users at 262K context. The author is soliciting feedback on context length trade-offs and access methods (API vs WebUI).
More from Infra
- Parlor: Open-source, on-device real-time multimodal AI similar to GPT-Live — tom_doerr · 2026-08-18
- Opinion: transformer will eventually be replaced — can Nvidia disrupt itself and stay ahead? — yangyi · 2026-08-18
- File Systems Emerge as Core Paradigm for AI Data Interaction — blaizedsouza · 2026-08-18
- Merge Partners with Mastra to Provide Unified API Gateway for AI Agents — shensi · 2026-08-18
- Does High Concurrency Make MoE Serving Load Nearly All Weights Per Token? — LocalLLaMa_reader · 2026-08-18
- Running Qwen3.8-27B with 256k Context on a Single 16GB GPU: Full Guide — ndiphilone · 2026-08-18