New "Discovery Episode" Framework Measures AI Scientists by Real Research Cycles

量子位 · wechat · 2026-08-24

A joint paper by DeepPrinciple, Microsoft, and Stanford introduces the "discovery episode" evaluation framework, shifting from closed-book exams like HLE to assessing AI's full research cycle. The system tracks decision-making trajectories through hypothesis, design, execution, and interpretation, valuing failed data as core assets. DeepPrinciple's MIRA platform, a real-world application of this framework, ranks first on the ResearchClawBenchmark and ScienceAgentArena, demonstrating its ability to close the loop between dry and wet labs.

Original post →

More from Research

Research channel →