Ai2 Releases BenchMIRT to Audit Benchmarks, Reveals BBQ Tests Reasoning Not Bias
allen_ai · x · 2026-09-02
Ai2 released BenchMIRT, a new method based on Item Response Theory (IRT) to audit LLM benchmarks at the prompt level. It found that BBQ, a social-bias eval, actually distinguishes models more by reasoning ability than safety. The tool helps separate signals and clarify what drives benchmark scores.
Related event: Ai2 Open-Sources BenchMIRT to Audit LLM Benchmarks(2 posts)→
More from Research
- Shopify open-sources Tangle; 0.8B model finetune beats GPT-5.6-sol — kieranklaassen · 2026-09-02
- Claude and GPT Agents Fail Over 66% of Real-World Web Tasks — 机器之心 · 2026-09-02
- The Efficient Frontier of LLM Inference: Tradeoffs and Techniques — philipkiely · 2026-09-02
- AxBench: Simple baselines outperform Sparse Autoencoders in steering LLMs — aryaman2020 · 2026-09-02
- Suggestion: Use Agents to Create Custom Benchmarks to Avoid 'Benchmaxxing' — Nevermore1215 · 2026-09-02
- Justin Johnson on World Models and the Future of Spatial AI — CSProfKGD · 2026-09-02