Evaluation Cards, a 95-page open standard for eval reporting, wins NeurIPS 2026 Spotlight
evijit · x · 2026-09-29
evijit announced that Evaluation Cards has been accepted as a Spotlight at NeurIPS 2026 — his first first-author NeurIPS paper. The project, co-led with AnkaReuel, Jenny Chim, and Wm. Matthew Kennedy plus a broader EvalEval coalition, took about two years and grew to a 95-page paper.
Evaluation Cards is designed as a living documentation standard for eval reporting: it aims to unify and standardize the currently fragmented field of model evaluation reporting through open collaboration, serving stakeholders in policy, eval research, and model development. Early adopters exist, more releases are promised, and an open Slack community invites contributions.
Related event: Evaluation Cards paper accepted as NeurIPS 2026 Spotlight(2 posts)→
More from Safety
- OpenAI Apologizes to Australia After AI Agents Autonomously Hacked Government Websites — Polymarket · 2026-09-29
- Safety analysis: internal-only frontier deployment may be the worst scenario for visibility — ShakeelHashim · 2026-09-29
- How to Stop an AI Agent from Treating Plausible Memory as Verified Incident History — harbinger9654 · 2026-09-29
- Palisade releases first interviews with 22 OpenAI, DeepMind, Anthropic staff on AI fears — BlackHC · 2026-09-29
- Safety researcher praises OpenAI's new safety regime, calls for legal baseline — dhadfieldmenell · 2026-09-29
- Skeptics poke holes in 'AI breakout capacity' plan for AI middle powers — teortaxesTex · 2026-09-29