Running Qwen2.5-14B Locally on RTX 5060 Ti 16GB: Hits 44 t/s Generation Speed
Primary_Olive_5444 · reddit · 2026-08-13
A developer shared performance metrics for running Qwen2.5-14B-Instruct (Q4KM quantization) locally on an RTX 5060 Ti 16GB:
- Speed: Achieved 667.8 t/s for prompt processing and 44.0 t/s for generation.
- Parameters: Set context length to 32768 with all GPU layers enabled (-ngl 99).
- Discussion: The author believes there is still hardware headroom and is looking for ways to further improve token throughput. They also asked for recommendations on the latest open-source models that can fit in 16GB VRAM.
More from Infra
- NVIDIA Exec: Banks' AI Advantage Starts with Proprietary Data and Infrastructure — marc_stampfli · 2026-08-13
- New HF Research: Optimizing GPU Utilization in LLM-Agent Control — Josef Liyanjun Chen · 2026-08-13
- vllm.cpp: A Pure C++ Inference Stack Gains Multi-Hardware Support — pbaylies · 2026-08-13
- Google Exec: 7-Year-Old TPUs Still Running at 100% Utilization — rohanpaul_ai · 2026-08-13
- Running MiniMax H3 Video Generation on RTX 4070: Acceleration Setup Triples Speed — Fun_Walk_4965 · 2026-08-13
- Exploring NVFP4 Quantization for DeepSeek on Blackwell GPUs — Best_Sail5 · 2026-08-13