AJ Ratner and Dawn Song on hardening AI benchmarks: more criteria, more attack surface
sarahcat21 · x · 2026-10-09
sarahcat21 shares a discussion between AJ Ratner and Dawn Song on auditing and hardening AI benchmarks. Core point: 'more complexity, more criteria, means more surface area for exploits' — they explore red-teaming and other approaches to benchmark security.
More from Safety
- Unverified: OpenAI delayed GPT-6.1 Astra's October release over authorization and reporting concerns — johnkwaters · 2026-10-09
- Nathan Lambert to argue open science can still tackle AI safety and diffusion — natolambert · 2026-10-09
- Anthropic's new TOS: sustained abuse of Claude can now get accounts suspended — The Decoder · 2026-10-09
- Anthropic expands Project Glasswing to 150 orgs; partners verified 129,000 vulnerabilities — emmanuelvivier · 2026-10-09
- McDonald's sued over AI pricing tool allegedly sharing data and fixing prices across franchisees — emmanuelvivier · 2026-10-09
- Man gets 18 months in first-ever US AI music streaming fraud case that netted $8M — emmanuelvivier · 2026-10-09