Snorkel AI Adds Bonus Livestream on How Benchmarks Are (and Aren't) Used in Model Development

dlwh · x · 2026-10-09

Snorkel AI added a bonus livestream on how benchmarks are and aren't used in model development today, featuring Braden Hancock (Laude Institute) and dlwh (Marin/Open Athena), covering Terminal-Bench Science, JudgmentBench, and SlopCodeBench with Steven Dillmann, Russell Yang, GOrlanski, and vincentsunnchen.

Original post →

More from Research

Research channel →