RTX 5090 Config for Qwen 27B Local LLM: Params & Tips
Rollingsound514 · reddit · 2026-08-27
A user shares their specific models.ini configuration for running Qwen 2.5 27B GGUF on an RTX 5090 with 128GB RAM. The setup includes Q5KXL quantization, Flash Attention, a 262K context window, and Q80 KV cache. The user seeks advice on trading quantization for context length to improve coding accuracy.
More from Infra
- 320B Model Runs on Mac: OrcaSAQ Quantization Brings GLM-5.3 to Apple Silicon — alejandroll10 · 2026-08-27
- GLM 5.3 Flash hits RunInfra day one: 254 tok/s, 1M context, $0.10/1M input — alejandroll10 · 2026-08-27
- Google Colossus: SSD performance at HDD prices — cis_female · 2026-08-27
- MiniMax H3 Max sets new Pareto Frontier for video gen, 50x faster. — Recoil42 · 2026-08-27
- NVIDIA Unveils NVHBM: 30% Higher Bandwidth and 15% Lower Power — zephyr_z9 · 2026-08-27
- Nikesh Shora predicts NeoCloud valuations will drop below current levels in 2 years — salgar · 2026-08-27