BenchmarkList Tracks 2.4k+ AI Benchmarks
davidthesong · reddit · 2026-07-16
BenchmarkList is a public platform aggregating 2.4k+ AI benchmarks, models, and capability results. It offers continuously updated evaluation streams, model comparisons, and in-depth research pages like the "AI progress on human work" dashboard and visualizations.\n\nThe author notes that current AI capability evaluations are scattered across papers, GitHub, model cards, websites, and tweets. This unified entry point aims to help researchers systematically understand what models "can and cannot do."
Related event: BenchmarkList Launches with Over 2,400 AI Benchmarks(4 posts)→
More from Research
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11