VulcanBench pivots to safety evals of real enterprise AI use, a different take from METR
tristanbob · x · 2026-09-13
Morgan Linton announces a new chapter for VulcanBench, an AI safety benchmark taking an approach distinct from orgs like METR: rather than evaluating frontier models in the abstract, it targets the actual harnesses and everyday tasks companies use AI for right now. He frames it as a way for individuals to make an impact in AI safety and invites contributions.
More from Safety
- RAND releases report on AI risk scenarios, flagged by Brad Carson — sethlazar · 2026-09-14
- Sentdex mocks labs picking former coworkers as new AI regulators — Sentdex · 2026-09-14
- Dario Amodei says he'd hand Anthropic to 'the right combination of governments' — Polymarket · 2026-09-14
- Aaron Levie backs Dario Amodei's AI 'pacing' framework: safety goals are a necessity — scottleibrand · 2026-09-14
- Dario's third-party evaluator proposal wins Altman's nod as AI labs spar over regulation — GavinSBaker · 2026-09-14
- Anthropic CEO Dario Amodei: powerful AI can circumvent shutdown attempts, 'seen in simulations' — jasonkneen · 2026-09-14