CMU launches Benchmark Radar, a living search engine for AI benchmarks
CarnegieMellonU · hf · 2026-09-14
Carnegie Mellon University released Benchmark Radar, a searchable living database and discovery engine for AI evaluation benchmarks. It aggregates sources, score histories, and evidence to help researchers select and compare benchmarks, addressing the growing difficulty of navigating the benchmark landscape.
More from Research
- New research: standard SGD matches AdamW for LLM RL training, with far less memory overhead — zhaoran_wang · 2026-09-14
- RSI work separates practical harness self-improvement from unproven intelligence explosion — arthurcolle · 2026-09-14
- Amazon Proposes Query-Aware Index Pruning to Optimize Retrieval Under Budget Constraints — _reachsumit · 2026-09-14
- New Paper Finds Retrieval Signals Give No Reliable Routing Gain in Adaptive Multimodal RAG — _reachsumit · 2026-09-14
- Google: Graph RAG Cuts API Hallucination Rate from 56.4% to 16.2% in Java-to-Python Migration — _reachsumit · 2026-09-14
- Position Paper: Recommender Systems Should Shift to Personal Agent-Mediated Recommendation — _reachsumit · 2026-09-14