Probability-Only Model Jev Beats or Matches GPT-5.6-luna on 42 of 49 Tasks at 4x Speed, 1/4 Cost

LowNefariousness9966 · reddit · 2026-09-19

A Redditor benchmarked TypeSafe's Jev — a model that generates no text, only probabilities for typed questions (yes/no, multiple choice, ratings) — against gpt-5.6-luna across 49 tasks and 8,200 items (SST-2, MMLU, ARC, MS MARCO, XNLI, SQuAD, Banking77, etc.).

Key findings

Where it lost

The author notes contamination risk, that the baseline was held to no/low reasoning, and that sub-0.05 differences are noise. Jev doesn't replace your LLM — it replaces the small classifier, reranker, and relevance calls around it. Full benchmark harness and raw results are open-sourced.

Related event: Probability-Only Jev Model Matches GPT-5.6 in 42 of 49 Tasks(3 posts)→

Original post →

More from Models

Models channel →