Snorkel AI Adds Bonus Livestream on How Benchmarks Are (and Aren't) Used in Model Development
dlwh · x · 2026-10-09
Snorkel AI added a bonus livestream on how benchmarks are and aren't used in model development today, featuring Braden Hancock (Laude Institute) and dlwh (Marin/Open Athena), covering Terminal-Bench Science, JudgmentBench, and SlopCodeBench with Steven Dillmann, Russell Yang, GOrlanski, and vincentsunnchen.
More from Research
- Could Looped Models Resist Distillation Attacks by Reasoning in Latent Space? — moyix · 2026-10-09
- LLM-as-a-Verifier: Weaker Model Verifies Stronger One, Hits 69.2% SOTA on Terminal-Bench 4 — Azaliamirh · 2026-10-09
- Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6% — Salesforce · 2026-10-09
- Claude claims discovery of new binary red dwarf pair ~500 light-years away — nitarshan · 2026-10-09
- Prompt Tuning Is Forgotten Lore — Are We Massively Underusing Finetuned Tokens? — cephaloform · 2026-10-09
- Ricardo Baeza-Yates Lecture: When Will ML Evaluation Stop Fooling Itself? — PolarBearby · 2026-10-09