RTX 3090 runs Qwen2.5-72B at 2,000 tok/s prefill
iamMess · reddit · 2026-09-01
A developer optimized Qwen2.5-72B (referred to as Qwen3.8-27B) on an RTX 3090, achieving 2,000 tokens/s prefill and 132 tokens/s decode. The core improvement is a custom kernel that matches fp32 quality with 0.99997 similarity at int8. The author believes decode speed is currently maxed out and has provided a GitHub repo for testing.
More from Infra
- Energy, Grid, Cooling First: Why AI's Real Competitive Stack Starts Below Compute — ingliguori · 2026-09-01
- Comet releases Opik, an open-source LLM observability tool for debugging and monitoring — dl_weekly · 2026-09-01
- VMware Expert: Private AI Offers Economic and Compliance Edge — DavidLinthicum · 2026-09-01
- Obscura: A Lightweight Rust Browser Built Exclusively for AI Agents — Shruti_0810 · 2026-09-01
- Space Data Centers Cost 20x More to Launch Than to Build on Earth — aronchick · 2026-09-01
- Switching to AMD 9060XT: Compatibility and speed for image gen — Tayunskapon · 2026-09-01