Apodex launches TRACES, a benchmark grading AI's reasoning trajectories on unsolved science problems
rohanpaul_ai · x · 2026-09-04
- Apodex released TRACES, a benchmark built for hard scientific problems where the correct answer may not exist yet — addressing the gap that most benchmarks assume an answer exists.
- How it works: problems become executable environments where a solver can use data and tools, receive feedback, and revise its approach, leaving a full trajectory. Trajectories are scored across six capabilities: Tools, Repair, Alternatives, Coherence, Evidence, and Scope, with a hidden verifier separately evaluating the outcome — so failures can be traced to wrong tool choice or failed recovery.
- The technical report, led by chief scientist @profshengwang, involved 10 STEM PhDs over two months surveying 561 industries to curate 423 real high-stakes problems spanning biomedicine, clinical translation, and frontier-model engineering. Both problem and solver submissions are open.
More from Research
- Researcher: the best papers answer nothing and raise more questions — yacinelearning · 2026-09-04
- IBM releases DRACO: dynamic rubrics give per-step credit assignment for verifier-free agent RL — ibm-research · 2026-09-04
- The Last Translation Benchmark debuts with multimodal examples that break leading translation models — Vilém Zouhar · 2026-09-04
- Reed-Solomon List Decoding Breakthrough Resolves Major Open Problem in Coding Theory — AlexKontorovich · 2026-09-04
- AI autoresearchers help crack 30-year-old coding theory problem, soundness up to 68.02 bits — BenBlaiszik · 2026-09-04
- Redditor fine-tunes SDXL on 60 childhood photos to simulate memory recall — uisato · 2026-09-04