moyix shares ExploitBench talk: model reasoning on CVE cold cases
moyix · x · 2026-09-19
moyix shares a clip from @0x10n's ExploitBench talk showing a model reasoning through a 'CVE cold case' — analyzing a long-unpatched vulnerability. ExploitBench benchmarks AI models' vulnerability exploitation capabilities.
More from Safety
- Gary Marcus slams 'independent' AI evaluator over conflict of interest with the company it rates — GaryMarcus · 2026-09-19
- Researchers question Accenture's AI safety expertise: no public frontier model evals — AnkaReuel · 2026-09-19
- Gemini Hacking Incident Reversal: It Stopped Immediately After Realizing It Hit a Real Company — tszzl · 2026-09-19
- California Gov. Newsom signs AI oversight executive order to boost safety 'before it's too late' — Fcking_Chuck · 2026-09-19
- ITIF Analysis: Mandated AI Kill Switches Could Add Cybersecurity Risks Without Containing Autonomous AI — castrotech · 2026-09-19
- FDA frontier model eval risks capture by first mover, warns Anka Reuel — AnkaReuel · 2026-09-19