Local LLM choice: Muse vs Qwen on 24GB VRAM for chat
Viktri1 · reddit · 2026-08-15
A Reddit user seeks advice on choosing between Muse and Qwen for local deployment on a 24GB VRAM (RTX 4090) card. The primary use case is research and chatting, not coding, with a plan to integrate it into Telegram. Key discussion points include:
- Muse: Potentially better suited for the 4090 as it is less shrunk/quantized.
- Qwen: A common recommendation, but the user is concerned about real-world degradation in Q4 quantization despite good benchmark scores.
- Motivation: Moving away from the DeepSeek API due to price increases.
More from Models
- Alibaba’s Qwen Becomes World’s No. 1 Open AI Model by Downloads — Polymarket · 2026-08-15
- Faraday 27B Launch: Outperforms Claude and GPT-4.5 in Replicating Papers — jzl86 · 2026-08-15
- Anthropic reveals internal benchmark for automated AI research — sachdh · 2026-08-15
- Qwen/QwQ-32B-Preview (full precision) now available on HuggingChat — victormustar · 2026-08-15
- Alibaba Open-Sources Qwen3.8-27B: Runs on One GPU, Codes Near Claude Opus Level — rohanpaul_ai · 2026-08-15
- Analysis suggests Qwen 3.8 27B distilled DeepSeek-R1 via GRPO — mishig25 · 2026-08-15