AI2 releases BenchMIRT to audit LLM benchmarks
Hugging Face Blog · rss · 2026-09-02
AI2 introduced BenchMIRT, a new method for auditing LLM benchmarks question by question. It reveals which capabilities benchmarks actually measure and helps researchers build smaller, more focused, and easier-to-interpret evaluations.
More from Research
- Legacy DS to LLM: Use banking77 dataset, SFT, and GRPO — JoshPurtell · 2026-09-02
- Shopify open-sources Tangle; 0.8B model finetune beats GPT-5.6-sol — kieranklaassen · 2026-09-02
- Claude and GPT Agents Fail Over 66% of Real-World Web Tasks — 机器之心 · 2026-09-02
- The Efficient Frontier of LLM Inference: Tradeoffs and Techniques — philipkiely · 2026-09-02
- AxBench: Simple baselines outperform Sparse Autoencoders in steering LLMs — aryaman2020 · 2026-09-02
- Suggestion: Use Agents to Create Custom Benchmarks to Avoid 'Benchmaxxing' — Nevermore1215 · 2026-09-02