Local model benchmarks: Opus 4.8 vs. DeepSeek, Qwen, and GLM across agent & coding tests

perelmanych · reddit · 2026-09-01

A summary table comparing benchmark scores of current popular local models like DeepSeek-V4-Flash, Qwen3.8, and GLM-5.3 against Opus 4.8. The data covers Agentic capabilities, Coding, General abilities, and Multimodal tasks. DeepSeek and Qwen show strong performance in various tests, while Opus 4.8 remains competitive. Specific benchmarks include Terminal Bench, SWE-bench, and GPQA Diamond.

Original post →

More from Models

Models channel →