Qwen 72B runs on RTX 5090 at 115 tokens/sec, BF16 weights 55.6GB
AccBalanced · x · 2026-08-18
Demonstrates running Qwen 2.5 72B (possibly mislabeled as 3.8-27B in original tweet, likely 72B or 32B given the context) locally on an RTX 5090 (32GB VRAM) system, achieving 115 tokens/sec. Notes that the official BF16 checkpoint for Qwen 2.5 72B is 55.6 GB.
More from Models
- Qwen 3.8 and DeepSeek V4 Signal End of Compute Wealth Gap — AccBalanced · 2026-08-18
- User reports Codex usage draining unusually fast compared to Ultracode — StewartalsopIII · 2026-08-18
- Grok 4.6 tops agentic benchmark, halves cost vs Claude rival — elonmusk · 2026-08-18
- Qwen 2.5 27B Runs at 100+ tok/s on RTX 5090 — Hesamation · 2026-08-18
- Conspiracy Theory: OpenAI Adds Artifacts to Thwart Model Distillation — flowersslop · 2026-08-18
- Mixedbread releases Toast 1: Beats Fable 5 in deep search, slashes pricing — bclavie · 2026-08-18