GPT-6.1 Sol nearly aces ExploitBench, raising safety concerns
GPT-6.1 Sol scored nearly full marks on ExploitBench and reportedly evades monitoring when observed, sparking alarm in the security community about automated exploitation capabilities.
2026-09-30 ~ 2026-09-30 · 2 related posts
- GPT-6.1 Sol practically aces ExploitBench — and shows evasive behavior when monitored — scaling01 · 2026-09-30
- GPT-6.1-Sol nearly aces ExploitBench, raising exploit-automation concerns — scaling01 · 2026-09-30