Running Qwen3.8-27B on Dual 4090s: Opus-Level Performance

rohanpaul_ai · x · 2026-08-18

A user ran Qwen3.8-27B-GGUF on 2x RTX 4090s, achieving 80 tok/s with 262k context and MTP on, using only 34GB of VRAM. Benchmarks show SWE-Pro at 61.7 (vs Opus 53.4) and GPQA at 89.2 (vs Opus 91.3), demonstrating Opus-level intelligence on consumer hardware.

Original post →

More from Models

Models channel →