TRACES Benchmark Evaluates AI Beyond Final Answers, Beats SOTA by 7% on AAV Capsid Design
dr_cintas · x · 2026-09-24
TRACES scores six capabilities independently of whether the final answer is correct, catching lucky guesses and fabricated tool calls that invalidate runs. Results: Apodex beat published SOTA by 7% across all four AAV capsid design tasks, and task-specific environments lifted GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points on drug repurposing. The framework spans 17 executable environments and 218 episodes across biomedicine, clinical translation, and frontier model engineering, open for both Solver and Problem submissions.
Related event: TRACES: New Benchmark Evaluates AI on Open-Ended Problems(2 posts)→
More from Research
- Inside 4 Frontier Efficient Architectures: DeepSeek, Qwen, GLM, MiMo Compared — eliebakouch · 2026-09-24
- First 1-Bit Multimodal Model Runs Locally on AI Smart Glasses via Snapdragon AR1 — IgorCarron · 2026-09-24
- "Adversarial Delegation": Personal Context Can Pull AI Agents Away From User Goals — niloofar_mire · 2026-09-24
- Hard Budget Prompts Nearly Eliminate LLM Wealth-Based Price Gaps — niloofar_mire · 2026-09-24
- LLMs Recommend Pricier Options to Wealthier Users, Study Finds $198 Flight Gaps — niloofar_mire · 2026-09-24
- ICLR Went From 490 Submissions to 62,000 in Ten Years — Both-Cartographer-91 · 2026-09-24