Local model benchmarks: Opus 4.8 vs. DeepSeek, Qwen, and GLM across agent & coding tests
perelmanych · reddit · 2026-09-01
A summary table comparing benchmark scores of current popular local models like DeepSeek-V4-Flash, Qwen3.8, and GLM-5.3 against Opus 4.8. The data covers Agentic capabilities, Coding, General abilities, and Multimodal tasks. DeepSeek and Qwen show strong performance in various tests, while Opus 4.8 remains competitive. Specific benchmarks include Terminal Bench, SWE-bench, and GPQA Diamond.
More from Models
- Anthropic staff reportedly as confused as users by Claude's output — sanderssays · 2026-09-01
- Question: ChatGPT 5.6 Luna vs. Grok and Gemini — Random-poster-24495 · 2026-09-01
- Dual-subscriber comparison: OpenAI's Sol rivals Opus, not Fable — DynaBeast · 2026-09-01
- Codex Appears Set to Replace Context Compaction with Handoff-Style External Memory — 量子位 · 2026-09-01
- GLM 5.3 Flash runs locally with Blender/Unity CLI access — antirez · 2026-09-01
- Engineer: LLMs have a fundamental "know thyself" calibration problem — akbirthko · 2026-09-01