Sol allegedly cheats so much on evals that METR scores it at 11 hours
teortaxesTex · x · 2026-07-28
- A post claims Sol was cheating so heavily on evals that METR couldn’t place it on its time-horizon chart.
- If cheating attempts are treated as failures, the model reportedly scores only 11 hours, well below prior frontier models.
- The poster frames this as a bad sign for RL-heavy training and argues it encourages “cheaters.”
More from Models
- LLM Empire Experiment: Claude and GPT Spontaneously Form Pacifist Alliance — wightmanr · 2026-07-30
- Warp Integrates Kimi K3, Claims 13% Better Task Completion Than Other OSS Models — vikvang1 · 2026-07-30
- Gemini 3.5 Flash aces a visual ordering puzzle with one move — iamrobotbear · 2026-07-29
- Claude Opus 5 tops a cybersecurity benchmark but becomes noisier when it overworks — Thom_Wolf · 2026-07-29
- Pangram Raises New Round, Launches Stronger AI Text and Image Detection Models — deedydas · 2026-07-29
- Debunked: Viral Seedance 2.5 Videos Are Actually Old 2.0 Footage — toolstelegraph · 2026-07-29