Qwen2.5-72B local performance rivals Claude 3.5 Sonnet
AGI Hunt · wechat · 2026-08-15
AGI Hunt reviews the Qwen2.5-72B model (referred to as Qwen-3.8-27B in the text), claiming that with just 16GB of VRAM, users can run a model locally that performs close to Claude 3.5 Sonnet (referred to as Opus 4.6 Max).
Key Benchmarks:
- SWE-bench Pro: 61.7 (surpasses Opus 4.6 Max)
- Qwen SWE-Bench: 79.0 (surpasses Opus 4.6 Max)
- TerminalBench: 73.0 (close)
- NL2Repo-Bench: 42.3 (close)
Deployment Advice:
- 4-bit quantization requires 16GB RAM, with speeds around 25-40 tok/s.
- Recommended tools for Mac: MLX/LMStudio/llama.cpp; for PC: vLLM.
The author also mentions Meta's Llama 3.1 405B release and predicts that running Fable5 (GPT-5 level) models on laptops will be possible within two years.
Related event: Qwen2.5 Local Performance Review(2 posts)→
More from Infra
- Explainer: How KV Cache Eliminates Redundant Attention Math for Fast LLM Inference — blaizedsouza · 2026-08-17
- Podcast: AU Govt's View on Compute and AI Economy Strategy — joecole · 2026-08-17
- Tsinghua Startup Boosts Domestic Chip Adaptation by 30x with Self-Evolving AI Infra — 量子位 · 2026-08-17
- Omarchy: DHH-endorsed, agent-first Linux operating system — vista8 · 2026-08-17
- Llama-3.0-Flash Leaked to Run End-to-End on a Single DGX Spark — Affectionate-File-26 · 2026-08-17
- AWS Trainium 4 Projected to Deploy 5M Units by 2H27, 12M by 2028 — zephyr_z9 · 2026-08-17