New Benchmark Live: Independent Measure of AI Usage Based on 24k Real Conversations
sanmikoyejo · x · 2026-08-18
A new independent, public measure of how people actually use AI has launched, built from 24,521 consented conversations across 52 models. Led by Anka Reuel and Shayne Redford, the project was featured on the front page of the Washington Post and in a long piece in MIT Technology Review.
More from Research
- Artificial Analysis Launches Search Index Benchmarking Agent Search APIs — ArtificialAnlys · 2026-08-19
- Artificial Analysis Launches Search API Index: Parallel, Exa Lead — ArtificialAnlys · 2026-08-19
- ICML Paper: Measuring LLM conceptual consistency to reduce contradictions in agents — soumitrashukla9 · 2026-08-19
- CFP: Continual World Models Workshop @ NeurIPS 2026 — cindy_x_wu · 2026-08-19
- DataSmith auto agent outperforms model architecture tweaks using 200x fewer tokens via data interventions — josh_wills · 2026-08-19
- Pander Score evaluates sycophancy in AI models — eli_lifland · 2026-08-19