edith-1 monitors agent traces, beats Sonnet-5.5 on balanced accuracy at 1/274 the cost
xennygrimmato_ · x · 2026-10-09
teknopawn introduced edith-1, a decision model for online trace monitoring trained to catch agent failures, reward hacking, and problems with tasks, graders and environments. On their eval it beats Sonnet-5.5 in Claude Code on balanced accuracy at an estimated 1/274 of the cost. It serves as a probabilistic filter; flagged runs go to a slower, more expensive agent judge for evidence-grounded failure analysis across task, trajectory, verifier and execution environment.
More from coding & agent
- Agent Buddy: an open-source desk buddy that shows your AI coding agents' status — DanWahlin · 2026-10-09
- Developer's Grok Bot now natively integrates with X, no API or credits needed — daniel_mac8 · 2026-10-09
- AutoScientist's two-agent checklist loop auto-audits every training example — sarahookr · 2026-10-09
- Meta's KernelAgent uses multi-agent orchestration for 2.02x Triton kernel speedups — PyTorch · 2026-10-09
- Claude recovers lost 2019 build paths to recompile MakerDAO's DAI to an exact bytecode match — devanshmehta · 2026-10-09
- 6 models tested on real MCP servers: Opus 5.5 leads, open models cost 87% less per attempt — shensi · 2026-10-09