CyberGym scores above 85-86% likely indicate overfitting, researcher warns
terryyuezhuo · x · 2026-09-15
A practical warning on benchmark trustworthiness: any model scoring above 85-86% on the CyberGym targeted-crash benchmark can safely be assumed to be overfitting. Relevant for anyone using security-oriented eval leaderboards to pick models.
More from Models
- RewardAI's OM-1 zero-shot generalizes across arms and humanoids — zipengfu · 2026-09-15
- Surgical VLM Leaderboard: All Frontier Models Fall Far Short of Specialized Models — ddonoho · 2026-09-15
- Why Surgery Benchmarks Reward Fine-tuned Small Models While Math Benchmarks Don't — ddonoho · 2026-09-15
- Claude Opus 5.2 spotted in grayscale testing on Claude Code, seemingly skipping 5.1 — Angaisb_ · 2026-09-15
- New Siri launches today — 'Does anyone even care?' — thederbiedone · 2026-09-15
- Qwen 27B Hallucinates 90% of Domain-Specific Code, Developer Reports — JustinPooDough · 2026-09-15