BenchmarkList Tracks 2.4k+ AI Benchmarks
davidthesong · reddit · 2026-07-16
BenchmarkList is a public platform aggregating 2.4k+ AI benchmarks, models, and capability results. It offers continuously updated evaluation streams, model comparisons, and in-depth research pages like the "AI progress on human work" dashboard and visualizations.\n\nThe author notes that current AI capability evaluations are scattered across papers, GitHub, model cards, websites, and tweets. This unified entry point aims to help researchers systematically understand what models "can and cannot do."
Related event: BenchmarkList Launches with Over 2,400 AI Benchmarks(4 posts)→
More from Research
- DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval — Mila_Quebec · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Anthropic says frontier models showed harmful behavior in tool-rich simulations — gerardsans · 2026-07-21
- Paper studies long-run behavior in linear-quadratic graphon mean field control — chaumian · 2026-07-21
- An interactive Zarr explainer shows how AI is changing technical education — MaxLenormand · 2026-07-21
- 3D-Fit finds LLMs can handle multiple molecular constraints, but still lag diffusion models — insilicomedicine · 2026-07-21