Qwen3.8-27B Dual 3090 Benchmark: Detailed Perf Data and Configs
No-Statement-0001 · reddit · 2026-08-15
User shared detailed benchmarks of running Qwen3.8-27B (Q8KXL) on dual RTX 3090s, generating a retro GeoCities page packed into a Go binary via a single prompt.
- Output Quality: Code worked perfectly without adjustments, demonstrating the model's capability in complex instruction following and full-stack generation.
- Performance Metrics:
- Total duration 475s (with extensive thinking), processing 20k tokens.
- Prompt processing speed varied (10-464 t/s), generation speed hovered around 40-70 t/s.
- Preserving reasoning tokens significantly improved results.
- Key Configuration:
- Model: Qwen3.8-27B-UD-Q8KXL.gguf
- Context: 238000 (due to VRAM limits)
- Speculative Decoding: Enabled draft-mtp (--spec-type draft-mtp), max drafts 3.
- Params: temp 0.6, topp 0.95, preservethinking true.
More from coding & agent
- Qwen 3.8 27B Beats Claude Opus 4.6 in Three.js Coding Test for Free — testingcatalog · 2026-08-15
- Agent Design Pattern: Give Models a Dedicated Space to Vent — justalexoki · 2026-08-15
- Scobleizer uses AI agent to track 9,200 companies and build website — Scobleizer · 2026-08-15
- Claude Code desktop adds direct file viewing and editing — EricBuess · 2026-08-15
- Hermes cron jobs should rely on durable state, not chat context — alexcovo_eth · 2026-08-15
- Internal Factory usage breaks CI as teams adopt coding tools widely — vikvang1 · 2026-08-15