Evals are vanishing from model cards since ChatGPT went mainstream
evijit · x · 2026-09-14
The author points to their team's prior work examining which evaluations stopped being disclosed in model/system cards after ChatGPT's mainstream commercialization — and the picture is bleak. Full analysis is in their ICML paper covering 186 first-party reports and 248 third-party evaluation sources.
Related event: ICML Paper Finds AI Social Impact Evaluations Shrinking(3 posts)→
More from Research
- whitetree: dynamic exact kNN without rebuilds, 40-300x faster than sklearn BallTree — monononon34 · 2026-09-14
- Expressive power isn't what gradient descent finds: why RNN cells excel at state-tracking generalization — mike64_t · 2026-09-14
- Google AI x Econ team finds field evidence that prior expertise drives learning from AI-assisted work — soumitrashukla9 · 2026-09-14
- Graph theory from hometown streets: navigation, genome assembly and Hamiltonian paths — TivadarDanka · 2026-09-14
- Weekend hack: Project Titania reimplements Qwen3-0.6B from transformer to GPU ISA simulator — generativist · 2026-09-14
- Paperclip indexes 7.5M+ papers as an agent-native filesystem, with MCP server support — james_y_zou · 2026-09-14