brier: calibrated typed decisions for open models without fine-tuning — 1B models jump from 16% to 80%+

Spirited_Conference9 · reddit · 2026-10-03

brier is an open-source tool that gives self-hosted open models the 'decision model' capability of TypeSafe's closed Jev API: typed questions (which team? refund? urgency 1-5?) answered with calibrated probabilities in one pass, no text to parse.

Raw open-model token probabilities are position-dependent and badly calibrated — reversing options flips Falcon3-1B-Base's answer 99% of the time on 20-way questions. brier reads the answer distribution directly from next-token log-probs (no generation, no fine-tuning) and corrects it in three levels: L0 averages all option rotations plus an optional prior (zero labels); L1 temperature-scales with 50+ labels; L2 fits a tiny linear head on a hidden layer in seconds on CPU.

Results on banking20 (20 intents, 2,603 test messages): zero labels lift Falcon3-1B-Base from 16% to 60% and LFM2-1.2B from 20% to 62%; with 300 labels every model across six families hits 80-86% (Phi-4-mini best at 0.86); Qwen3-1.7B's ECE drops from 0.357 to 0.035.

15 models tested including MoEs and chat-template-less base models, plus a brier check preflight command, a free Colab tutorial, and a 10-minute add-your-model PR flow. Limits: one benchmark so far, transformers-only (no vLLM), L0 costs K forward passes.

Original post →

More from Research

Research channel →