Probability-Only Model Jev Beats or Matches GPT-5.6-luna on 42 of 49 Tasks at 4x Speed, 1/4 Cost

LowNefariousness9966 · reddit · 2026-09-19

A Redditor benchmarked TypeSafe's Jev — an odd model that generates no text, only probabilities in response to typed questions (yes/no, multiple choice, ratings) — against gpt-5.6-luna across 49 tasks and 8,200 items (SST-2, MMLU, ARC, MS MARCO, XNLI, SQuAD, Banking77, etc.) with identical prompt wording.

Key findings

Where it lost

The author notes contamination risk (the synthetic math set is the only control), that the baseline was deliberately held to no/low reasoning, and that with 200 items per task differences under 0.05 are noise. Jev can't write text, call tools, or explain itself — it doesn't replace your LLM, but the small classifier, reranker, and relevance calls around it. Task builders, runner, scorer and raw per-item results are open-sourced (stdlib-only Python).

Related event: Probability-Only Jev Model Matches GPT-5.6 in 42 of 49 Tasks(3 posts)→

Original post →

More from Models

Models channel →