4x RTX 3090 advice: keep Qwen 27B Q8 or switch to Qwen Next Flash
Zyj · reddit · 2026-09-19
The author is upgrading a Threadripper Pro 5955WX / 128GB DDR4 server from two to four RTX 3090s, currently serving a few users with Qwen 3.8 27B Q8 on vLLM.
The dilemma: keep running the 27B Q8 and use spare VRAM for smaller parallel models, or switch to Qwen 3.8 Next Flash at Q4/Q5 (which fits but may not be universally better). The Club 3090 multi-GPU doc is called out as horribly outdated. They're asking the community for recommendations.
More from Infra
- SpaceX targets first orbital AI compute satellites next year, sun as the power plant — XFreeze · 2026-09-19
- Open models hit record 78.4% of token volume on Vercel AI Gateway; Kimi, DeepSeek, Z.ai spend beats OpenAI — charles_irl · 2026-09-19
- Awesome MCP Servers: a 95k-star index of every production MCP server — mdancho84 · 2026-09-19
- Ollama hits 181k stars: run open-source LLMs locally in one command — mdancho84 · 2026-09-19
- X open-source algorithm update: weights unchanged, video checks start at 64 likes — ThePeterMick · 2026-09-19
- antirez: API pricing is Monopoly money — the only real metric is joules, and we can't see them — mitsuhiko · 2026-09-19