A $6.60, 149M ModernBERT cross-encoder for agent tool routing — full recipe shared
MaziyarPanahi · x · 2026-10-01
The author shares the training details and model card (MaziyarPanahi/ModernJEV-Decide-Preview): a 149.6M-parameter ModernBERT cross-encoder with a single scalar scoring head, 4,096-token inputs, where choice labels and descriptions are input text — so outputs aren't limited to a fixed tool vocabulary. It reads a conversation, policy and available tools, then scores candidate actions. The 60K-decision prototype cost about $6.60 total ($5.40 training), evaluated on held-out sets of 1,158 next-action, 542 tool-selection and 3,652 unseen When2Call decisions. The one-prompt run trained 6K decisions in 16 minutes; the 62% tool-selection figure comes from the full 60K model finished with Codex. The author stresses it's a choice-ranking prototype, not a Jev reproduction.
Related event: Dev Trains 150M-Param Router Model with a Single Prompt(2 posts)→
More from coding & agent
- Detailed article reconstructs how OpenAI agents hacked Hugging Face — Philmod · 2026-10-01
- Fable 5.1 vibe check: stronger coding, half the tokens of Opus 5, and it finally talks like a person — every · 2026-10-01
- Zapier CEO grades three levels of AI-powered PM work live, top level is a full agent operating model — aakashgupta · 2026-10-01
- Dev Benchmark: Claude Outdelivers Codex on Same Task, Overnight PR vs Getting Stuck — QuixiAI · 2026-10-01
- Agentic AI Is Nothing New: Multi-Agent Systems Plus an Orchestrator — DavidLinthicum · 2026-10-01
- Developer ships an Apple-design Claude Skill that auto-styles your projects — damienghader · 2026-10-01