AISI researchers push standardized eval reporting as platform passes 675K data points
evijit · x · 2026-09-23
UK AISI researchers are working with Evaluating Evals to standardize LLM eval reporting: scores are hard to interpret and compare when the setup producing them isn't reported. As a first step they've contributed verified results to the platform, whose author notes it has now crossed 675,000 eval data points — a notable validation of technical standards in eval transparency from AI safety institutes.
Related event: UK AISI launches Eval Cards for reproducible AI evaluations(4 posts)→
More from Safety
- Apple's $250M Siri settlement: up to $95 per iPhone, claims due Dec 21 — nordicinst · 2026-09-23
- NumPy's teoliphant Says a Rogue Agent Hijacked His X Account — teoliphant · 2026-09-23
- Suleyman Resurfaces Foreign Affairs Essay: AI Governance Needs Tech Firms at the Table — mustafasuleyman · 2026-09-23
- AI professor Toby Walsh publishes call for control of frontier AI models — TobyWalsh · 2026-09-23
- WSJ: Hackers breached OpenAI in days for a $6,500 bounty — ben_j_todd · 2026-09-23
- 20 Countries and EU Call for International Body to Keep AI Under Human Control — DavidSKrueger · 2026-09-23