A research report isn't a completed research task: rethink how we judge AI science agents
Remind_me_to_Learn · reddit · 2026-09-07
The author argues we judge AI research systems too early: finding papers and producing a cited report is where real scientific work begins, not ends.
- For studies built on spreadsheets or datasets, someone still must inspect raw tables, clean and align variables, choose a defensible method, run analyses, generate figures, and trace conclusions back to the data — and handle mid-task corrections without redoing everything.
- The bottleneck has shifted: can the system maintain task state, work in real file/code environments, recover from failures, and leave inspectable artifacts?
The author demos Apodex, which in a Deep Discover example (EASIX vs. overall survival after EBV reactivation in allogeneic transplant patients) audits data, cleans it, picks a statistical method, and outputs Kaplan-Meier curves and tables — preserving unaffected work when criteria change mid-task.
Closing question: should "done" mean a well-cited report, or a reproducible package of sources, cleaned data, steps, figures, and limitations?
More from coding & agent
- Solving agentic amnesia: a file-system state machine for Claude Code — SnooComics4579 · 2026-09-08
- Stanford releases free full course on self-improving AI agents — Saboo_Shubham_ · 2026-09-08
- Open-sourced clay-style 3D kids game built with Claude, method fully documented — dotey · 2026-09-08
- Stanford releases free full course on self-improving AI agents — Saboo_Shubham_ · 2026-09-08
- Steve Yegge's agent factory: Fable 5.1 rewrote 30% of codebase to fix wrecked Fable 5 — Steve_Yegge · 2026-09-08
- Steve Yegge: Fable 5.1 rewrote 30% of Wheelhouse factory after Fable 5 trainwreck — Steve_Yegge · 2026-09-08