TRACES Benchmark Evaluates AI Beyond Final Answers, Beats SOTA by 7% on AAV Capsid Design

dr_cintas · x · 2026-09-24

TRACES scores six capabilities independently of whether the final answer is correct, catching lucky guesses and fabricated tool calls that invalidate runs. Results: Apodex beat published SOTA by 7% across all four AAV capsid design tasks, and task-specific environments lifted GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points on drug repurposing. The framework spans 17 executable environments and 218 episodes across biomedicine, clinical translation, and frontier model engineering, open for both Solver and Problem submissions.

Related event: TRACES: New Benchmark Evaluates AI on Open-Ended Problems(2 posts)→

Original post →

More from Research

Research channel →