Qwen3.8-Max Ties Opus 5 in Cybersecurity Benchmark, Lacks Consistency

teortaxesTex · x · 2026-08-05

A security firm benchmarked Qwen3.8-Max on finding recent CVEs. Across three runs, Qwen identified 26 out of 32 CVEs (81.25% pass@3 recall), tying with Opus 5 for first place at half the cost.

However, Qwen showed significant inconsistency: only 10 CVEs were found in all three passes, compared to 19 for Opus. Commenters noted that while the base model is strong, its post-training appears half-done, leading to unstable performance.

Original post →

More from Models

Models channel →