MLST: Training AI to Replicate Papers — Faraday Beats Codex and Claude on Held-Out Tasks
Machine Learning Street Talk · rss · 2026-09-12
Machine Learning Street Talk hosts Edward Hughes, Chief Scientist and co-founder of Inherent, on whether machines can learn the judgment that separates plausible-looking results from faithful experiments.
Core arguments
- Creativity is not optimization: Move 37 was innovative, not creative — the field, not the individual, decides what counts as a discovery. The missing AI capability is choosing which questions are worth asking.
- The argument weaves in Csikszentmihalyi's creativity psychology, David Deutsch's hard-to-vary explanations, exaptation and open-endedness: deceptive goals and imperfect world models are the point, not the problem.
The paper
- Replica: a task space built by redacting figures from real papers for held-out replication.
- Faraday: a 27B-parameter model trained to steer a frontier coding agent; it beats Codex, Claude and GLM 5.2 on held-out replications.
Further topics include from-replication-to-innovation (how the Transformer happened), whether the AI scientist can Goodhart the judge, 8xB300 training runs, the RL crisis of per-turn credit in GRPO, recursive companies crossing a phase transition, and what replaces OKRs. Full timestamps and references available in the episode.
More from Research
- Schmidhuber revisits history: 1971 deep network modeled the British economy — SchmidhuberAI · 2026-09-20
- Sort samples by loss and check the head: a practical labeling-error hack — antoine_chaffin · 2026-09-20
- Qwen researcher: human annotation may be noisier than auto-labeling — antoine_chaffin · 2026-09-20
- Retrieval eval datasets are error-prone; LLM relabeling proposed as fix — CShorten30 · 2026-09-20
- PhyFilter from Beihang and NTU lets sim-trained quadrupeds walk real terrain, in npj Robotics — jiqizhixin · 2026-09-20
- ICLR submissions triple to 60,000 abstracts; reviewing would take 250 human-years — gleech · 2026-09-20