FULL STORY

Apodex Launches TRACES, First Benchmark for Discovery AI

Apodex, founded by Chen Tianqiao, released TRACES, a benchmark claiming to measure AI's ability to make genuine scientific discoveries rather than retrieve known answers.

2026-08-20 ~ 2026-09-11 · 5 episodes · 25 posts

Episode 1 · Apodex Launches TRACES, First Benchmark for Discoverative AI (2026-08-20, 8 posts)

On August 20, Apodex AI released TRACES, billed as the first benchmark for discoverative AI. Unlike existing benchmarks that test retrieval of known answers, TRACES targets research questions with no answer key, evaluating whether systems can process evidence, validate hypotheses, and reach verifiable conclusions — closer to real scientific discovery. Several commenters see it as pushing evaluation from "getting it right" toward "defensible discovery processes."

Confirmed

  • TRACES was released by Apodex AI as the first benchmark for discoverative intelligence
  • The release includes a definition of "discoverative intelligence"
  • Evaluation shifts from final-answer correctness to the inquiry process: tools used, errors caught, underlying evidence, and whether the process is rigorous, traceable, and recoverable from mistakes (per @alifcoder, @rohanpaulai)
  • Per @Divpradeep, Apodex's business converts research questions into executable environments for AI solvers to test hypotheses, and claims its systems are trained accordingly

Why it matters

  • Commenters such as @alifcoder and @SimonShaoleiDu argue that benchmarks built on known-answer tasks no longer differentiate model capability; TRACES poses a harder, more substantive question: whether AI can make defensible scientific discoveries when the answer is unknown
  • If widely adopted, it could set a more meaningful standard for scientific AI

Episode 2 · Apodex Launches TRACES Benchmark for AI Scientific Discovery (2026-08-22, 2 posts)

Apodex has released TRACES, billed as the world's first benchmark for evaluating "discovery-oriented AI." Unlike traditional retrieval-based tests, it assesses an AI's ability to use tools, gather evidence, test hypotheses, and draw evidence-based conclusions without a ready-made answer key.

Episode 3 · Apodex Unveils TRACES Benchmark to Measure AI's Ability to Discover (2026-08-25, 3 posts)

Apodex, founded by Tianqiao Chen, has released TRACES, the first benchmark for "Discoverative AI." Unlike traditional tests that rely on known answer keys, TRACES measures AI's ability to make scientific discoveries about the unknown.

Episode 4 · Apodex Releases TRACES: A Benchmark for AI Scientific Discovery on Open Questions (2026-09-04, 9 posts)

On September 4, Apodex (founded by Tianqiao Chen) released TRACES, a new benchmark positioning itself under a "Discoverative AI" paradigm. Endorsed by researchers including dair-ai and Omar Sanseviero (@omarsar0), it is the first systematic attempt to evaluate research agents making credible discoveries on open scientific questions whose answers are not yet confirmed.

Confirmed

  • TRACES does not test whether models answer known questions; it measures whether they can rigorously and evidence-based advance unknown problems. Its name is an acronym of six capabilities: the first three assess the exploration process, the last three assess conclusion quality—Coherence (consistency across long investigations), Evidence (every claim traceable to real sources), and Scope (the solver states the boundaries of its conclusions).
  • The 423 problems were hand-crafted by 10 PhDs, characterized by "the correct answer may not yet exist."
  • Evaluation uses dual verifiers: an outcome verifier checks only results against hidden ground truth, while a process verifier reviews logged operations, errors, and corrections without seeing result scores.
  • Per @SucceededMind, TRACES claims "the era of static benchmarks is over": it evaluates the agent's full execution loop—tool selection and strict tool execution, dynamic error repair, and long-context handling—rather than comparing final outputs to hidden answers.
  • Process scores also drive a repair loop: per Apodex's report relayed by omarsar0, flagged runs generate repair notes, and solvers retry without seeing answers; re-running failed trajectories yields an average gain of 0.155. In an AAV capsid design case, the process-verifier-driven repair loop surpassed SOTA.
  • The benchmark targets a blind spot: HLE, FrontierMath, MMLU, and BrowseComp are hard but assume answers exist; models that ace them can still stall on genuinely open research.

Why it matters

  • TRACES fills the scientific-discovery dimension traditional benchmarks cannot measure, shifting evaluation from "answering known questions" to "rigor in exploring the unknown."
  • The process-verifier-driven repair loop demonstrates that process scores can do more than grade, potentially shaping the training and evaluation of next-generation research AI.

Episode 5 · Apodex unveils TRACES benchmark for discovery-oriented AI (2026-09-11, 3 posts)

Apodex, founded by Chen Tianqiao, released TRACES, a benchmark with a live leaderboard that evaluates 'discoverative AI' on 423 open scientific questions, scoring the reasoning process across six dimensions rather than checking final answers.