RTX 5090's 32GB VRAM vs 256k Context: Is It Enough for Local Qwen 27B Agentic Coding?
lots_of_puppies · reddit · 2026-09-25
A Reddit user runs Qwen 27B locally on a 128GB M5 Max, but prompt processing and output slow down badly as context grows, and even the faster Qwen Next struggles for heavy agentic coding. He's eyeing a Ryzen 9 9950X3D + RTX 5090 build hoping for 3000+ PP and 80-150 tok/s, but the 5090's 32GB VRAM means dropping to Q4 quantization and possibly losing full 262k context. He asks whether the build is worth it.
More from Infra
- Goldman Sachs hikes AI power forecasts: 2030 data center capacity raised to 217GW — McDonaghMatthew · 2026-09-25
- Puro-2B: an open recipe trains a Qwen2-1.5B-beating LLM on RTX 5090s for just $4.4K — IgorCarron · 2026-09-25
- YC-backed Prism launches agent-optimized inference cloud, serving DeepSeek V4.1 Flash at 547 tok/s — ycombinator · 2026-09-25
- Going local-first: 2x 5060 Ti at ~100 tok/s plus a 96GB M5 Ultra, betting on a 3-year inference floor — No-Name-Person111 · 2026-09-25
- Google's Suncatcher project aims to put AI data centers in orbit powered by solar energy — The Decoder · 2026-09-25
- Setting up a DGX Spark local inference cluster with 5 role-based Hermes agents — natesiggard · 2026-09-25