Médula open-sources data on coordinating parallel Claude Code agents: Jev decider escalates 61% of real writes
jokiruiz · reddit · 2026-10-01
The author of Médula, an MIT-licensed kernel coordinating multiple Claude Code agents on one repo, published surprising decision-layer data. A PreToolUse hook intercepts writes; a fast decider (Jev) scores collision probability, with a slow path (Sonnet→Opus) for ambiguous cases.
Lab results: Jev passed 37/37 tests at $0.29 decision cost per run vs Haiku's $0.55; median latency 279ms vs 1338ms; $0.00006 per decision. Calibration on 100 pairs: Jev 92.5%, Haiku 91.5%, Sonnet 98.5%.
The lab-to-real gap: only 13% of calibration decisions escalated, but 61% of real write requests did—making the decision layer cost half of an all-Sonnet setup rather than a tiny fraction.
Eval lesson: careless labels scored Jev 68%; careful labels scored the same answers 87%. Caveats: 1-5 runs per mode, thresholds fit on measured data, labels written by Claude-family models. Repo: github.com/JoaquinRuiz/medula.
More from coding & agent
- Developer wires Jev into Codex to auto-route each task to the right model — aziz4ai · 2026-10-01
- Workaround for giving coding agents multiple machines: SSH chaining — pvncher · 2026-10-01
- cliffhanger: A Stop hook that makes Claude Code finish its task list unattended — Arthur122103 · 2026-10-01
- Polyphonic lets you carry your web agent memory into ChatGPT, Codex and Claude — RileyRalmuto · 2026-10-01
- Designing MCP tool output: raw rows, summaries, or answers with confidence scores — No-Plant-5234 · 2026-10-01
- Security checklist before wiring AI coding agents to a production database — idontlikepeople26 · 2026-10-01