Local LLM Speed Bottlenecks: RTX 4090 vs. 5090 Performance Analysis
Viktri1 · reddit · 2026-08-19
The user analyzes the factors affecting the speed of local LLMs, including token efficiency, memory bandwidth (RAM vs. VRAM), and the prefill phase impacted by GPU compute and prompt length. Running Qwen 3.8 27B on an RTX 4090 yields 1200 tok/s prefill and 80 tok/s generation. The user questions whether upgrading to an RTX 5090 is necessary for higher speeds or if a 48GB modified 4090 suffices. The conclusion is that VRAM capacity primarily affects model quantization and cache size, while generation speed (tokens/s) is compute-bound, making the 5090 the superior choice for raw speed despite its high cost. The user also notes that while the Muse model is faster, its coding capabilities are significantly inferior to Qwen's.
More from Infra
- AMD posts async RL walkthrough on MI355X and benchmarks vs B300 — AnushElangovan · 2026-08-19
- Andrej Karpathy releases llm.c: Train LLMs in raw C/CUDA — goyalshaliniuk · 2026-08-19
- Anthropic's $50B Buildout Shows Financing Is Not the Short-Term Compute Bottleneck — FinanceYF5 · 2026-08-19
- Epoch AI: funding won't bottleneck frontier compute; model to scale past 20GW — FinanceYF5 · 2026-08-19
- Five project companies issued $15.18B in debt for 1.43GW of data centers — FinanceYF5 · 2026-08-19
- Anthropic leveraged under $9B revenue into nearly $50B AI infrastructure — FinanceYF5 · 2026-08-19