FAccT Researcher: Safety Evals Ignore Rich Measurement Work Beyond Benchmarks
evijit · x · 2026-09-13
Researcher evijit argues today's AI safety discourse over-relies on benchmarks while ignoring richer measurement traditions from the FAccT community. Collaborators struggle with evals that are expensive and slow on university hardware, and niche lived experiences are missing from mainstream system cards. He calls for more thoughtful access to compute for diverse evaluation efforts.
More from Safety
- Dario calls to pace the frontier as Anthropic pledges permanent third-party eval access — niloofar_mire · 2026-09-13
- METR hiring AI safety evaluators with salaries up to $687K, matching top labs — chrisrohlf · 2026-09-13
- Regulation is coming for open and closed AI models alike — the question is proactive or reactive — benjamin_warner · 2026-09-13
- Lawyers weigh in: defendants would win cases over open source model licenses — markjeffrey · 2026-09-13
- Not regulatory capture but ass-covering: labs fear public backlash more than open source — willcb · 2026-09-13
- AI safety critic calls e/acc's stance a fundamental mistake in retweeted thread — NathanpmYoung · 2026-09-13