Running Qwen3.8-27B on RTX 5060 Ti 16GB: IQ4 Quantization and Full Feature Benchmark

Tema_Art_7777 · reddit · 2026-08-22

The author tested Qwen3.8-27B on an RTX 5060 Ti 16GB to validate 64K context, vision, and agentic tool use on a single card. Comparing jpetrina IQ4XS-pure, Unsloth UD-IQ4XS, and Q80, results show Unsloth UD-IQ4XS with MTP-1 maintains 45 tok/s at 64K context with perplexity very close to Q8. Vision (F16 projector) works but strains VRAM when combined with max context, suggesting profile separation. In BFCL tool-use benchmarks, the 27B models significantly outperformed Qwen3.5-9B in multi-turn workflows, handling complex scenarios like dependent calls and prompt injection resistance.

Original post →

More from coding & agent

coding & agent channel →