Visualizing Benchmarks: Qwen 3.8-Max Outperforms Opus 4.8 Across Multiple Metrics

deliprao · x · 2026-08-13

To better visualize LLM benchmark data, Deliprao recreated a performance comparison chart.

He highlighted the specific evaluation tasks where Qwen 3.8-Max numerically beat Opus 4.8 in light red. The redrawn chart clearly shows that Opus 4.8's capabilities are largely subsumed or surpassed by Qwen 3.8-Max across most benchmarks. The author cautiously noted that "beat" strictly means numerically exceeding the score, and some pairs may lack statistical significance.

Original post →

More from Models

Models channel →