Review: Running Qwen3.8-27B at 128K Context on a Single 32GB GPU
WSTangoDelta · reddit · 2026-08-15
A user reported running Qwen3.8-27B Q6K GGUF with 128K context on a single 32GB R9700 GPU via llama.cpp. The model was tasked with auditing a real, legacy swimming pool controller codebase. It spent 21 minutes tracing architecture and logs to produce a 30KB report, demonstrating strong understanding of the messy codebase, with agentic behavior holding up well compared to previous versions.
More from Models
- Questioning how to detect Gemini degradation — ngxson · 2026-08-15
- Unsloth Releases Qwen3.8-27B-NVFP4 — TheLocalLab · 2026-08-15
- User Feedback: Opus 5 Seems Too Histrionic — curious_vii · 2026-08-15
- OpenAI's Unreleased 'Astra' Model Solves 10 Math Problems — thursdai_pod · 2026-08-15
- Long Reasoning Task Consumes 22k Tokens for 3k Output — generativist · 2026-08-15
- Users Report Codex Asking for Permission More Often — GabGarrett · 2026-08-15