A research report isn't a completed research task: rethink how we judge AI science agents

Remind_me_to_Learn · reddit · 2026-09-07

The author argues we judge AI research systems too early: finding papers and producing a cited report is where real scientific work begins, not ends.

The author demos Apodex, which in a Deep Discover example (EASIX vs. overall survival after EBV reactivation in allogeneic transplant patients) audits data, cleans it, picks a statistical method, and outputs Kaplan-Meier curves and tables — preserving unaffected work when criteria change mid-task.

Closing question: should "done" mean a well-cited report, or a reproducible package of sources, cleaned data, steps, figures, and limitations?

Original post →

More from coding & agent

coding & agent channel →