Local LLM Test: Qwen2.5 72B Hits 100 tok/s on RTX 4090

julianharris · x · 2026-08-25

A user tested Qwen2.5 72B (q4 quantized) with the Unsloth framework on an RTX 4090, achieving 100 tok/s inference speed. Although the context window was compressed to around 66k, the model successfully handled a complex Rust project with GUI requirements, compacting context twice before delivering a working result.

Original post →

More from coding & agent

coding & agent channel →