Local LLM Test: Qwen2.5 72B Hits 100 tok/s on RTX 4090
julianharris · x · 2026-08-25
A user tested Qwen2.5 72B (q4 quantized) with the Unsloth framework on an RTX 4090, achieving 100 tok/s inference speed. Although the context window was compressed to around 66k, the model successfully handled a complex Rust project with GUI requirements, compacting context twice before delivering a working result.
More from coding & agent
- Zero-LLM MCP Tool Diagnoses Apollo GraphQL Cache Corruption — dev_nihar · 2026-08-25
- VectorSmith: Controlled Vector DB Access for Agents via YAML — dontgimmehope · 2026-08-25
- 100 AI Personas Simulate Reddit: They Form Factions and Hold Grudges — mrjeeves · 2026-08-25
- VectorSmith: Define Vector DB Tools in YAML for MCP — dontgimmehope · 2026-08-25
- Open Source Project 'Munder Difflin': A Multi-Agent Coding Harness — Saboo_Shubham_ · 2026-08-25
- SaaS Future Prediction: API-First Companies Embracing MCP Will Replace Closed Agent Vendors — garrytan · 2026-08-25