BenchmarkList Tracks 2.4k+ AI Benchmarks

davidthesong · reddit · 2026-07-16

BenchmarkList is a public platform aggregating 2.4k+ AI benchmarks, models, and capability results. It offers continuously updated evaluation streams, model comparisons, and in-depth research pages like the "AI progress on human work" dashboard and visualizations.\n\nThe author notes that current AI capability evaluations are scattered across papers, GitHub, model cards, websites, and tweets. This unified entry point aims to help researchers systematically understand what models "can and cannot do."

Related event: BenchmarkList Launches with Over 2,400 AI Benchmarks(4 posts)→

Original post →

More from Research

Research channel →