Ornith 9B runs coding agents for 3.5 hours on 16GB GPU

CrowKing63 · reddit · 2026-08-21

Finding Qwen 3.8 27B slow on a 16GB AMD GPU, the author tested Ornith 1.5 9B (Q6K, 256K context) with Q8 KV cache.

Results showed 950 tok/s prompt eval and 36 tok/s generation. It ran continuously for almost 3.5 hours on a real coding agent task. The author shared specific launch commands, suggesting that optimizing a 9B model is more practical than chasing larger models on low-VRAM hardware.

Original post →

More from coding & agent

coding & agent channel →