FAccT Researcher: Safety Evals Ignore Rich Measurement Work Beyond Benchmarks

evijit · x · 2026-09-13

Researcher evijit argues today's AI safety discourse over-relies on benchmarks while ignoring richer measurement traditions from the FAccT community. Collaborators struggle with evals that are expensive and slow on university hardware, and niche lived experiences are missing from mainstream system cards. He calls for more thoughtful access to compute for diverse evaluation efforts.

Original post →

More from Safety

Safety channel →