Apodex's TRACES benchmark: 10 PhDs hand-built 423 open research problems

omarsar0 · x · 2026-09-04

Elvis Saravia walks through TRACES, a new benchmark from Apodex targeting a blind spot in current evals: HLE, FrontierMath, MMLU, and BrowseComp all come with answer keys, so a model can top all of them yet still stall on genuine open research. TRACES uses problems whose ground truth may take months or years to confirm. Ten STEM PhDs spent two months scouting 561 industries across 16 sectors to hand-build a registry of 423 high-value problems.

Related event: Apodex Launches TRACES, a Benchmark for AI Scientific Discovery on Open-Ended Questions(8 posts)→

Original post →

More from coding & agent

coding & agent channel →