Researcher Slams AI Eval Orgs for Lacking Basic Software Engineering Practices

evijit · x · 2026-08-09

AI researcher evijit expressed frustration on X regarding the poor software engineering practices prevalent among AI evaluation organizations. The author pointed out that while evals underpin most AI policy discussions, many organizations fail to maintain basic version control and reproducibility, severely undermining the reliability of these benchmarks.

Original post →

More from Research

Research channel →