Médula open-sources data on coordinating parallel Claude Code agents: Jev decider escalates 61% of real writes

jokiruiz · reddit · 2026-10-01

The author of Médula, an MIT-licensed kernel coordinating multiple Claude Code agents on one repo, published surprising decision-layer data. A PreToolUse hook intercepts writes; a fast decider (Jev) scores collision probability, with a slow path (Sonnet→Opus) for ambiguous cases.

Lab results: Jev passed 37/37 tests at $0.29 decision cost per run vs Haiku's $0.55; median latency 279ms vs 1338ms; $0.00006 per decision. Calibration on 100 pairs: Jev 92.5%, Haiku 91.5%, Sonnet 98.5%.

The lab-to-real gap: only 13% of calibration decisions escalated, but 61% of real write requests did—making the decision layer cost half of an all-Sonnet setup rather than a tiny fraction.

Eval lesson: careless labels scored Jev 68%; careful labels scored the same answers 87%. Caveats: 1-5 runs per mode, thresholds fit on measured data, labels written by Claude-family models. Repo: github.com/JoaquinRuiz/medula.

Original post →

More from coding & agent

coding & agent channel →