Apodex launches TRACES benchmark for AI scientific discovery
On September 4, Apodex (founded by Chen Tianqiao) released a new benchmark called TRACES, which drew attention after being recommended and introduced by several AI researchers, including dair-ai and Omar Sanseviero (@omarsar0).
Confirmed
- TRACES is positioned as a new paradigm of "Discoverative AI": rather than testing AI's ability to answer known questions, it measures whether models can rigorously and evidence-basedly push forward exploration of unknown problems.
- The name TRACES is itself an acronym for six capabilities covering the exploration dimensions needed for scientific discovery.
- The questions were hand-crafted by 10 PhDs, totaling 423 open research problems characterized by "correct answers that may not yet exist."
- Evaluation goes beyond final answers: it turns research questions into assessable tasks while also examining the model's reasoning trajectory.
- The benchmark targets a blind spot in existing evaluations: HLE, FrontierMath, MMLU, and BrowseComp are hard, but they all assume answers already exist—a model can ace them all yet still stall on genuinely open-ended research.
Why it matters
- TRACES is the first systematic benchmark for whether research agents can make credible discoveries on open scientific questions with unconfirmed answers, filling a scientific-exploration dimension that traditional benchmarks cannot measure.
- If adopted by the community, it could shift evaluation focus from "answering known questions correctly" to "rigor in exploring the unknown," shaping how the next generation of research AI is trained and evaluated
2026-09-04 ~ 2026-09-04 · 8 related posts
Primary sources
- Apodex launches TRACES, a benchmark measuring whether AI can rigorously explore the unknown — eyishazyer · 2026-09-04
- Apodex's TRACES benchmark: 10 PhDs hand-built 423 open research problems — omarsar0 · 2026-09-04
- TRACES' last three axes grade conclusions: coherence, evidence, scope — omarsar0 · 2026-09-04
- Apodex reports process-verifier repair loop beats SOTA in AAV capsid design (+0.155 over 434 retried runs) — omarsar0 · 2026-09-04
- Process scores drive a repair loop: failed trajectories gain 0.155 on rerun — omarsar0 · 2026-09-04
- [source] TRACES: new benchmark scores AI agents on unconfirmed scientific discoveries — dair_ai · 2026-09-04
- Apodex launches TRACES, a benchmark grading AI's reasoning trajectories on unsolved science problems — rohanpaul_ai · 2026-09-04
1 near-duplicate retellings: omarsar0