Running Qwen3.8-27B on Dual 4090s: Opus-Level Performance
rohanpaul_ai · x · 2026-08-18
A user ran Qwen3.8-27B-GGUF on 2x RTX 4090s, achieving 80 tok/s with 262k context and MTP on, using only 34GB of VRAM. Benchmarks show SWE-Pro at 61.7 (vs Opus 53.4) and GPQA at 89.2 (vs Opus 91.3), demonstrating Opus-level intelligence on consumer hardware.
More from Models
- Researcher: GPT 5.6 Sol Ultra Beats Pro for Long-Horizon Hard Problems — arankomatsuzaki · 2026-08-24
- Google Criticized: Gemini 3.7 Still Missing From Its Own Jules Agent a Week Later — brandon_galang · 2026-08-24
- Qwen 27B 3.8 low quantization tested: Q3 XXS works well locally — jeremyckahn · 2026-08-24
- Users notice significant quality shift in GPT-5.6 output — haider1 · 2026-08-24
- Ramp Stats: Anthropic Opus 4.8 and Sonnet 4.6 Lead Usage — vista8 · 2026-08-24
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24