EvaluatingEvals Grows from 100K to 500K Eval Results Since Beta Launch
evijit · x · 2026-09-13
EvaluatingEvals launched its beta with 100K eval results and has since grown to 500K. The project is expanding its list of covered evaluators and is publicly soliciting more data sources.
More from Research
- Decagon on GEPA-GAN: Simulated Users That Are Too Cooperative Are Skewing Agent Evals — kastnerkyle · 2026-09-13
- AVERI Paper on Frontier AI Auditing Backs Dario's Embedded Evaluator Commitment — Miles_Brundage · 2026-09-13
- AI math proofs won't kill understanding: post hoc exploration keeps mathematicians central — njyx · 2026-09-13
- Two high schoolers, heavily AI-assisted, solve a problem that stumped Fields Medalist June Huh — QuanquanGu · 2026-09-13
- MHA, MQA, GQA and MLA explained by what happens to the K/V cache during decoding — techNmak · 2026-09-13
- SAE features reveal an "Immersive Simulation Mode" inside LLMs during roleplay — sebkrier · 2026-09-13