Ornith 9B runs coding agents for 3.5 hours on 16GB GPU
CrowKing63 · reddit · 2026-08-21
Finding Qwen 3.8 27B slow on a 16GB AMD GPU, the author tested Ornith 1.5 9B (Q6K, 256K context) with Q8 KV cache.
Results showed 950 tok/s prompt eval and 36 tok/s generation. It ran continuously for almost 3.5 hours on a real coding agent task. The author shared specific launch commands, suggesting that optimizing a 9B model is more practical than chasing larger models on low-VRAM hardware.
More from coding & agent
- Real-world Ops: Managing the Extreme Overhead of Production Agentic Systems — zakelfassi · 2026-08-21
- Grok Build Update: Native API Integration, Custom Domains, GitHub Export — XFreeze · 2026-08-21
- Speed up agent swarms: use bv to analyze bead dependencies for better concurrency — doodlestein · 2026-08-21
- Investigating Overhead and Decision Fatigue in Managing Agents at Scale — zakelfassi · 2026-08-21
- Workflow Tip: Connecting ChatGPT Pro to GitHub Beats Using Codex Alone — jdjohnson · 2026-08-21
- AMD MI300x vs NVIDIA H100: Real-world agent coding benchmark — locker73 · 2026-08-21