Visualizing Benchmarks: Qwen 3.8-Max Outperforms Opus 4.8 Across Multiple Metrics
deliprao · x · 2026-08-13
To better visualize LLM benchmark data, Deliprao recreated a performance comparison chart.
He highlighted the specific evaluation tasks where Qwen 3.8-Max numerically beat Opus 4.8 in light red. The redrawn chart clearly shows that Opus 4.8's capabilities are largely subsumed or surpassed by Qwen 3.8-Max across most benchmarks. The author cautiously noted that "beat" strictly means numerically exceeding the score, and some pairs may lack statistical significance.
More from Models
- DeepSeek V4 Pro Update Suspectedly Pulled Amid Abnormal Benchmark Scores — op7418 · 2026-08-13
- Qwen3.8-27B Coming Soon: Release Set for August 14 on Hugging Face — huggingface · 2026-08-13
- DeepSeek-V4-Pro Leak: Nears GPT-5.6 in Coding Benchmark at 1/31st the Price — rohanpaul_ai · 2026-08-13
- DeepSeek V4 API Fingerprint Changes, Hinting at New Checkpoints — teortaxesTex · 2026-08-13
- Users Report Severe Model Degradation Across Google's APIs — EthanBeMe · 2026-08-13
- Qwen3.8-27B Model Surfaces on ModelScope Ahead of Hype — Ok-Shower7286 · 2026-08-13