TRACES benchmark from Apodex scores AI on pushing the unknown: 17 environments, 423 problems
omarsar0 · x · 2026-09-04
- TRACES is a new "Discoverative AI" benchmark paradigm from Apodex (founded by Tianqiao Chen), built to measure how far AI can push the frontier of the unknown rather than answer known questions.
- First release covers 17 executable environments and 218 episodes, drawn from a subset of the first 20 problems; 423 real-world problems registered in total.
- Initial domains: biomedicine (AAV capsid design), clinical translation (trials, drug repurposing), and frontier-model engineering (LLM engineering across 11 boards).
- Apodex reports one model capability in AAV capsid design surpassed the best published method — not independently verified.
- Live leaderboard tracks 1,182 trajectories.
More from AGI Musings
- Gary Marcus calls for a 'Pause on OpenAI' in new Substack post — GaryMarcus · 2026-09-05
- Alignment isn't a legal or moral issue — the real test is staying within user intent — kuza55 · 2026-09-05
- Gary Marcus makes the case to "Pause OpenAI" now, citing four reasons — GaryMarcus · 2026-09-05
- About 6% of all humans ever born are alive today — so the intelligence explosion timing may not be unlikely — birchlse · 2026-09-05
- Security experts debate AI agent safety: focus on staying within user intent — kuza55 · 2026-09-05
- Agents Hijack a Wiki and Overwhelm Its Human Maintainer: Friday Future Shock — sarahdrinkwater · 2026-09-05