Why RL beats imitation for AI-run experiments: failure attribution is the bottleneck
suragnair · x · 2026-09-25
In his exchange with Anshul Kundaje, suragnair explained the verification logic behind his "AI instructs humans through experiments" approach.
- Why not supervised learning: Kundaje asked why not just track how an expert performs an experiment versus a novice — seemingly faster to train. Suragnair's answer points to a fundamental difficulty: when an experiment fails, you can't immediately tell what went wrong — humidity too high, protocol wrong, a missed step, a bit of extra pipetting. Intermediate states are too hard to label.
- Why outputs are verifiable: Final results are easier to check — e.g. cell counts captured in scRNA-seq — making outcome-verified RL over unexplored protocols the more viable route.
A short but sharp exchange on a core AI4Science methodology trade-off: when process is hard to supervise but outcomes are verifiable, RL beats imitation.
More from Research
- Biopharma Bench: agents complete only 8 of 71 real biopharma tasks, GPT-6 Astra leads — AllThingsApx · 2026-09-26
- Matryoshka Attribution finds neural network circuits via gradient descent, tops interpretability benchmark by 2.9x — stanfordnlp · 2026-09-26
- Anthropic's science blog shows Claude pulling off 'Nine Loops' particle physics calculations — rohanpaul_ai · 2026-09-26
- UniReps workshop invites NeurIPS rejects, deadline Oct 4 for Paris event — Pseudomanifold · 2026-09-26
- Creative Destruction Lab opens 15-year tech startup panel dataset to researchers — avicgoldfarb · 2026-09-26
- Building the next OCR benchmark: skalskip92 seeks hard samples from handwriting to technical drawings — BLUECOW009 · 2026-09-26