Frontier Data Summit lineup: Chollet, Dawn Song headline batch of new agent benchmarks
StanfordAILab · x · 2026-10-06
Snorkel AI's invite-only Frontier Data Summit (Oct 8, San Francisco) revealed its agenda, unveiling a wave of new benchmarks: STELLA-Bench, JudgmentBench, CollusionBench, HalluWorld, AgentAbstain, HumanOversight Bench, Terminal-Bench-Science, and Train-to-Test (T²) scaling laws. Speakers include François Chollet (ARC Prize), Dawn Song (UC Berkeley), and Sanmi Koyejo (Stanford STAIR Lab). Three themes anchor the event: evaluating agents in realistic tool-rich environments, measuring long-horizon autonomy, and scoring complex multi-artifact outputs.
More from Research
- Is the brain a computer? Physical reservoir computing pokes holes in Turing-simulation argument — cephaloform · 2026-10-06
- Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test — alexisjross · 2026-10-06
- Embedding Every Font with Neural Networks Yields a Flower-Shaped Map of Google Fonts — Chroma-Crash · 2026-10-06
- Crawler Zoo Launches a Free Arena for Testing Local-Model Agents — Time_Instruction_955 · 2026-10-06
- Trained agentic context management: 8K-context small model matches GPT-5.4 at 1M on OOLONG — xennygrimmato_ · 2026-10-06
- User Sim Index is broken: trivial bot scores 95% across behavioral dims — ericzelikman · 2026-10-06