Benchmarking Qwen 3.8 27B on RTX 4090: Q5 vs Q6 in 3D Code Generation
source-drifter · reddit · 2026-08-15
A user benchmarked Qwen 3.8 27B (Q5-xl and Q6-k) models on an RTX 4090 (24GB) for generating a Three.js 3D arena game. At 'xhigh' reasoning settings, the Q6 model took 3.5 hours to generate 84k tokens (6 t/s), while the Q5 model finished 57k tokens in 13 minutes (72 t/s). Although Qwen 3.8 outperforms version 3.6 in quality, the time and token cost are 3-4 times higher at maximum settings. At 'medium' reasoning, the Q5-xl offers better results in under 2 minutes.
More from Models
- DeepSeek V4 Pro tops benchmarks with specific configs, rivaling GPT-5.6 and Claude — teortaxesTex · 2026-08-15
- Qwen 3.8 35BA3B model spotted in GitHub commit — BazzyIm · 2026-08-15
- DeepSWE benchmarks spark re-evaluation of Fable model performance — teortaxesTex · 2026-08-15
- GLM-5.3 Review: Matches GPT-4 Coding, and I Built 3 Plugins with It — 赛博禅心 · 2026-08-15
- Test shows Qwen3.8-27b water surface rendering beats Gemini Flash — pbaylies · 2026-08-15
- OpenAI makes GPT-5.6 Luna the default free ChatGPT model with unlimited chats — emmanuelvivier · 2026-08-15