Apodex's TRACES benchmark: 10 PhDs hand-built 423 open research problems
omarsar0 · x · 2026-09-04
Elvis Saravia walks through TRACES, a new benchmark from Apodex targeting a blind spot in current evals: HLE, FrontierMath, MMLU, and BrowseComp all come with answer keys, so a model can top all of them yet still stall on genuine open research. TRACES uses problems whose ground truth may take months or years to confirm. Ten STEM PhDs spent two months scouting 561 industries across 16 sectors to hand-build a registry of 423 high-value problems.
More from coding & agent
- Developer Hosts Projects in Google Antigravity and Pairs It With Claude — oilmutt · 2026-09-05
- Dev building a Rust SSR framework with 'ridiculous' hydration benchmarks, asks for contenders — mohamedmansour · 2026-09-05
- LangChain hiring a lead for SmithDB, its database built for massive agent trace storage — LangChain · 2026-09-05
- Deploying agents that touch honeypot boards is risky — case-by-case calls and in-sandbox escalation needed — voooooogel · 2026-09-05
- Astra's agent workflow: keep working, only block when the answer changes the outcome — HaktanSuren · 2026-09-05
- Knowledge worker seeks cheaper Claude alternatives with strong instruction-following and long context — SemiMagnum · 2026-09-05