Snorkel Scales Open Benchmarks Grants to $30M as Marin Lead Explains How Evals Guide Training
dlwh · x · 2026-10-10
David Hall, lead of the Marin project, shared how benchmarks are actually used in model development at the Frontier Data Summit:
- Benchmarks sit in a 2x2 grid: pretraining vs posttraining, internal development vs external communication; the famous "model card evals" are just one quadrant
- The best internal evals work at multiple scales, guiding scaling laws and giving signal weeks or months before models score non-zero on traditional benchmarks
- Pretraining relies mostly on corpus-derived metrics like loss/perplexity, since modern benchmarks target post-trained models that may score zero for months
Meanwhile Snorkel AI announced a 10x expansion of its Open Benchmarks Grants to $30M, having funded 25+ open benchmarks (Terminal-Bench, OSWorld, Agents' Last Exam), with a steering committee including Chris Ré, Karthik Narasimhan, and Ludwig Schmidt.
Related event: Snorkel Expands Open Benchmarks Grants to $30M, Forms Red Team(3 posts)→
More from Research
- EA-VAE paper in IEEE T-PAMI fixes systematic uncertainty failures in VAEs — enzoferrante · 2026-10-10
- OpenAI reportedly solved 92 of the 500 most important open math problems in one GitHub push — altryne · 2026-10-10
- Isola lays out the three main pushbacks to the PRH narrative his new paper answers — phillip_isola · 2026-10-10
- Isola's team shows a global orthogonal map aligns image and text embeddings without paired data — phillip_isola · 2026-10-10
- Paper decodes 315K encrypted reasoning blocks, recovers 367 PII and 182 credentials — DynamicWebPaige · 2026-10-10
- Chollet: science is the ultimate recursively self-improving system, yet progress stays linear — fchollet · 2026-10-10