edith-1 monitors agent traces, beats Sonnet-5.5 on balanced accuracy at 1/274 the cost

xennygrimmato_ · x · 2026-10-09

teknopawn introduced edith-1, a decision model for online trace monitoring trained to catch agent failures, reward hacking, and problems with tasks, graders and environments. On their eval it beats Sonnet-5.5 in Claude Code on balanced accuracy at an estimated 1/274 of the cost. It serves as a probabilistic filter; flagged runs go to a slower, more expensive agent judge for evidence-grounded failure analysis across task, trajectory, verifier and execution environment.

Original post →

More from coding & agent

coding & agent channel →