Qwen3.8-Max Ties Opus 5 in Cybersecurity Benchmark, Lacks Consistency
teortaxesTex · x · 2026-08-05
A security firm benchmarked Qwen3.8-Max on finding recent CVEs. Across three runs, Qwen identified 26 out of 32 CVEs (81.25% pass@3 recall), tying with Opus 5 for first place at half the cost.
However, Qwen showed significant inconsistency: only 10 CVEs were found in all three passes, compared to 19 for Opus. Commenters noted that while the base model is strong, its post-training appears half-done, leading to unstable performance.
More from Models
- LiquidAI Releases LFM2.5-2.6B Edge Model — LiquidAI · 2026-08-05
- shadcn: I'd Trade Benchmark Points for Half the Latency — shadcn · 2026-08-05
- Rumor: SSI Developing Online Learning, GPT-6 Expected This Month — flowersslop · 2026-08-05
- Optimized MiniMax H3 Demo Space Available on Hugging Face — mrfakename0 · 2026-08-05
- Pokee-Isaac 28B Launches: 10M-Token Context on a Single GPU — Kyrannio · 2026-08-05
- Asked the Same Geography Question 8 Times, AI Reached the Opposite Conclusion Every Time — Voxtante · 2026-08-05