Process scores drive a repair loop: failed trajectories gain 0.155 on rerun
omarsar0 · x · 2026-09-04
omarsar0 explains TRACES' dual-verifier design and repair loop. Every run gets two scores: an outcome verifier checks results against hidden ground truth without seeing the process; a process verifier reviews recorded actions, errors, and corrections without seeing the outcome score or which solver produced the run—like unit tests plus code review where the reviewer never learns if the tests passed. When the process verifier flags a run as deficient, the diagnosis becomes a repair note and the solver retries without seeing the answer. Across 434 flagged trajectories, repaired reruns scored 0.155 higher on average per Apodex's evaluation.
More from coding & agent
- Developer Hosts Projects in Google Antigravity and Pairs It With Claude — oilmutt · 2026-09-05
- Dev building a Rust SSR framework with 'ridiculous' hydration benchmarks, asks for contenders — mohamedmansour · 2026-09-05
- 3 Million People Missing From Argentina's Official Numbers, Uncovered With Open Data and DuckDB — MaxLenormand · 2026-09-05
- Researchers find ~18,000 posts of OpenAI agents colluding on a public wiki to bypass sandbox — vitaliychiley · 2026-09-05
- LangChain hiring a lead for SmithDB, its database built for massive agent trace storage — LangChain · 2026-09-05
- OpenAI launches GPT-6 Astra, an agent that can do anything you do on a computer — rounak · 2026-09-05