Local LLM user on RTX 5090 weighs Qwen 27B vs Flash Next for coding: is a bigger model worth it?
MasterNomie · reddit · 2026-10-11
A local LLM user with an RTX 5090 shared model-selection experience and asked the community for advice:
- Started with Qwen 3.8 27B on Ninfer at NVFP4; decent overall but often ignored instructions in AGENTS.md and agent skills
- Switched to Flash Next for better breadth of analysis and foresight; currently happy
- Found local models outperform ChatGPT Go and free-tier Claude/Gemini on health, fitness, and ergonomics topics
Questions posed: is the quality jump to bigger models noticeable? Is sacrificing speed worth it? Do smarter models save total time? Has anyone upgraded, regretted, and gone back down? VRAM limits on a 5090 make the next step costly.
More from coding & agent
- This founder runs his entire product distribution from an Obsidian vault pointed at Claude Code — EXM7777 · 2026-10-11
- RSIGym pattern: keep agents in lightweight CPU containers, offload training/inference/evals to services — SucceededMind · 2026-10-11
- One settings change turns Claude Code into a multi-model team: Opus plans, Sonnet codes, Haiku searches — chessbuzz · 2026-10-11
- Stanford paper: fix looping agents via harness edits, bad plans need weight training — rohanpaul_ai · 2026-10-11
- GhidraMCP updated for latest Ghidra with headless support: load plugin and start reversing — lauriewired · 2026-10-11
- 13-year gamedev builds Claude Code-driven 3D asset pipeline, barely writing code by hand — Embarrassed_Guide_80 · 2026-10-11