Braintrust adds Jev as a judge scorer: typed decisions at up to 193.6× speed and 444.6× lower cost
multiply_matrix · x · 2026-09-19
- Braintrust now supports TypeSafe's Jev as a judge scorer for evaluating agent responses.
- Jev handles narrow, typed decisions over unstructured data: your code supplies state and constrained-answer questions, and it returns answers with probability and confidence in parallel — no prose parsing needed. Example: a double-charge complaint classified as billing with 100% confidence.
- In Braintrust you can inspect Jev's answer, confidence, and probabilities, and trace calls via JavaScript/Python SDKs to refine scoring criteria.
- TypeSafe reports up to 193.6× faster execution and 444.6× lower cost in its own workflow evaluations; the post also shows how to break judgments like "ready to send" into categories or scored choices.
More from coding & agent
- Claude Code now supports AGENTS.md as fallback when CLAUDE.md is absent — jasonkneen · 2026-09-19
- Pydantic Creator Benchmarks Jev vs Sonnet: 2.6x Cheaper, 3.7x Faster in 17 Lines — samuelcolvin · 2026-09-19
- GPT-6 Astra clears Geometry Dash demon level Jumper with all 3 coins — imjustnewatai · 2026-09-19
- Greg Kamradt: Humans with AI still beat AI with AI on productivity — GregKamradt · 2026-09-19
- Veteran ML engineer: Jev may push agent tool-calling back to discriminative models — multiply_matrix · 2026-09-19
- Box CEO demos Jev for instant enterprise document triage at near-zero cost — multiply_matrix · 2026-09-19