Eve, a 0.6B open model, passes calibration: when it says 90%, believe it
colinmcnamara · x · 2026-09-24
One conclusion from the calibration thread: Eve, a 0.6B open model, narrowly passed — it reached 90% confidence on 115 of 1,599 questions, and only 12 of its 673 wrong answers claimed 90%, so its confidence statements are largely reliable. The thread also tested Jev, a hosted decision API: 96% right on the hard questions it kept, no GPU needed, but it is not deterministic and a third of its wrong answers on human labels still claimed 90%.
Related event: Calibration Tests: Small Models Often Claim 90% Confidence When Wrong(2 posts)→
More from coding & agent
- RL tutorial: using Jev as reward model lifts Qwen reward from 0.583 to 0.759 — sophiamyang · 2026-09-24
- Running Qwen locally with Hermes agent in 10 minutes on an old gaming laptop — markjeffrey · 2026-09-24
- The Jev + Claude Code combo: a 100ms, near-free decision brain that blocks destructive AI commands — PrajwalTomar_ · 2026-09-24
- TensorSharp Adds Image-Aware Local Jev API With 3.3x Faster Structured Reads, Zero Schema Failures — fuzhongkai · 2026-09-24
- Resend ships Stripe Projects integration as the most-requested email provider — jeff_weinstein · 2026-09-24
- Anthropic cut Opus 5.5 prices, then broke four things your agent depends on — rseroter · 2026-09-24