TRACES: new benchmark scores AI agents on unconfirmed scientific discoveries
dair_ai · x · 2026-09-04
dair-ai highlights a new benchmark from Apodex AI: TRACES scores AI systems on discoveries where the answer isn't yet confirmed.
Unlike benchmarks with ground-truth answers, it measures whether research agents can produce credible findings on open problems — a capability the authors argue is crucial for agents used in real research.
More from Research
- Anima Anandkumar's first podcast on Accelerated Understanding: 5T context and 4D physical AI — AnimaAnandkumar · 2026-09-05
- Uno pairs AR weights with diffusion LoRA adapters, beating EAGLE-3 and all diffusion LLMs — HongyiWang10 · 2026-09-05
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05
- Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals — a_karvonen · 2026-09-05
- Deep Learning Weekly #471: Claude Fable 5.1 launch, production-parity LLM evals, alignment paper — dl_weekly · 2026-09-05