RTX 3090 runs Qwen2.5-72B at 2,000 tok/s prefill

iamMess · reddit · 2026-09-01

A developer optimized Qwen2.5-72B (referred to as Qwen3.8-27B) on an RTX 3090, achieving 2,000 tokens/s prefill and 132 tokens/s decode. The core improvement is a custom kernel that matches fp32 quality with 0.99997 similarity at int8. The author believes decode speed is currently maxed out and has provided a GitHub repo for testing.

Original post →

More from Infra

Infra channel →