Ai2 Releases BenchMIRT to Audit Benchmarks, Reveals BBQ Tests Reasoning Not Bias

allen_ai · x · 2026-09-02

Ai2 released BenchMIRT, a new method based on Item Response Theory (IRT) to audit LLM benchmarks at the prompt level. It found that BBQ, a social-bias eval, actually distinguishes models more by reasoning ability than safety. The tool helps separate signals and clarify what drives benchmark scores.

Related event: Ai2 Open-Sources BenchMIRT to Audit LLM Benchmarks(2 posts)→

Original post →

More from Research

Research channel →