Frontier Data Summit showcases a dozen new AI benchmarks and a push to rethink evaluation science

sanmikoyejo · x · 2026-09-30

Snorkel AI shared a Frontier Data Summit conversation between Stanford STAIR Lab director Sanmi Koyejo and Terminal-Bench co-creator Alex Shaw on reimagining the science of AI measurement and evaluation. The invite-only one-day summit (Oct 8, San Francisco) unveiled over a dozen new benchmarks — STELLA-Bench, ApprenticeBench, JudgmentBench, CollusionBench, HalluWorld and more — covering agent evaluation in realistic tool-rich environments, autonomy horizons, and verifiable multi-artifact scoring. Speakers included François Chollet and Dawn Song.

Original post →

More from Research

Research channel →