VulcanBench pivots to safety evals of real enterprise AI use, a different take from METR

tristanbob · x · 2026-09-13

Morgan Linton announces a new chapter for VulcanBench, an AI safety benchmark taking an approach distinct from orgs like METR: rather than evaluating frontier models in the abstract, it targets the actual harnesses and everyday tasks companies use AI for right now. He frames it as a way for individuals to make an impact in AI safety and invites contributions.

Original post →

More from Safety

Safety channel →