Real dev work fully local: Ornith 1.5 35B on 8GB VRAM hits 30-40 tok/s with OpenCode
Sensitive_Song4219 · reddit · 2026-09-23
After weeks of a 'local-AI-only' regime, the author reports solid agent coding on an RTX 5060 Mobile (8GB VRAM): Ornith 1.5 35b-a3b runs at 30-40 tok/s generation. Key points:
- Model choice: Qwen 3.8 27b remains the local coding star, but Ornith 1.5 35b-a3b is a strong 30b-a3b alternative — extra training gives it a 'High reasoning'-like depth, whereas Qwen 3.6 35b-a3b MoE only has on/off thinking and tends to underthink
- Harness: Ornith struggled with file writes in Pi; OpenCode via llama.cpp fixed it at 11k base context, and OpenCode enables seamless mid-chat switching to cloud models
- Tests: a web-OS demo worked with two bug-fix turns; fully local router-log troubleshooting identified a WPS-related WLAN crash in 10 minutes and autonomously built a log-scraper/monitoring app
Verdict: agentic coding on modest hardware is already quite usable.
More from coding & agent
- Claude Code v2.1.280+ lets you switch Opus 5.5 effort mid-session without breaking prompt cache — kimmonismus · 2026-09-23
- Sentry founder David Cramer finds Claude Code's new sidebar confusing vs Codex — zeeg · 2026-09-23
- Mid-session effort switching on Opus 5.5 preserves prompt cache in Claude Code v2.1.280+ — lydiahallie · 2026-09-23
- fable-advisor v6: open-source Claude Code plugin teams up Opus 5.5, Fable 5.1, GPT-6 Sol and Luna — daniel_mac8 · 2026-09-23
- Open-Source fable-advisor: Fable 5.1 Architect Orchestrating GPT-6 and Opus — daniel_mac8 · 2026-09-23
- GPT-6 Sol vs Grok 4.7: one prompt, one HTML file, three Star Wars worlds each — rohanpaul_ai · 2026-09-23