Local LLM Review: Qwen-27B Replaces Cloud, Ornith-9B Handles Coding Tasks
Barni275 · reddit · 2026-08-31
A developer shared hands-on experience running open-source models locally for coding:
- Qwen-2.5-72B (labeled 27B/3.8 in post): Runs on an RTX 3090 (Q80 KV, 180k context) with 1000-1500 tps prefill and 45-60 tps generation. The experience is excellent, replacing cloud subscriptions from 5-6 months ago for major coding tasks.
- Ornith-1.5-9B: Runs on a home server with 8GB VRAM. Without benchmarks, it successfully handled simple Python refactoring tasks. The author seeks more real-world coding feedback on this model.
More from coding & agent
- Dev shares a dirt-cheap approach to visual diffs — zeeg · 2026-09-01
- Grok Bot automates Shopify updates and supplier coordination — billyjhowell · 2026-09-01
- Investor calls GrokBot the next ChatGPT moment: 3 minutes beats hours of work — 新智元 · 2026-09-01
- Grok Bot automates lost deal analysis by mining call and email threads — lennysan · 2026-09-01
- Design pattern: immutable agent artifact revisions behind a stable review URL — RocketSeven · 2026-09-01
- Building a long-term memory benchmark for agents: what to add? — True_Mongoose_7073 · 2026-09-01