Process scores drive a repair loop: failed trajectories gain 0.155 on rerun

omarsar0 · x · 2026-09-04

omarsar0 explains TRACES' dual-verifier design and repair loop. Every run gets two scores: an outcome verifier checks results against hidden ground truth without seeing the process; a process verifier reviews recorded actions, errors, and corrections without seeing the outcome score or which solver produced the run—like unit tests plus code review where the reviewer never learns if the tests passed. When the process verifier flags a run as deficient, the diagnosis becomes a repair note and the solver retries without seeing the answer. Across 434 flagged trajectories, repaired reruns scored 0.155 higher on average per Apodex's evaluation.

Related event: Apodex Launches TRACES, a Benchmark for AI Scientific Discovery on Open-Ended Questions(8 posts)→

Original post →

More from coding & agent

coding & agent channel →