Review: Running Qwen3.8-27B at 128K Context on a Single 32GB GPU

WSTangoDelta · reddit · 2026-08-15

A user reported running Qwen3.8-27B Q6K GGUF with 128K context on a single 32GB R9700 GPU via llama.cpp. The model was tasked with auditing a real, legacy swimming pool controller codebase. It spent 21 minutes tracing architecture and logs to produce a 30KB report, demonstrating strong understanding of the messy codebase, with agentic behavior holding up well compared to previous versions.

Original post →

More from Models

Models channel →