TRACES: A New Benchmark That Grades AI Problem-Solving Process, Not Just Correct Answers

dr_cintas · x · 2026-09-24

Nearly every LLM benchmark relies on a fixed answer key. Apodex's TRACES targets problems where the answer isn't known yet, and is the first to do three things: evaluate the whole system (model, harness, tools, memory, control policies) rather than just the model; grade six process capabilities via an HDS6 rubric — Tools, Repair, Alternatives, Coherence, Evidence, Scope — independent of whether the final answer is right, catching lucky guesses and fabricated tool calls; and open submissions to teams. As the author notes, this mirrors what peer reviewers actually argue about: the process, not the conclusion.

Related event: TRACES: New Benchmark Evaluates AI on Open-Ended Problems(2 posts)→

Original post →

More from coding & agent

coding & agent channel →