Jev skeptics: BGE-small + logistic regression hits 93.3% on Banking77 vs Jev's 83.2%
tiensss · reddit · 2026-09-24
A Reddit post argues that Jev, marketed as a new class of "System One" decision model, is mostly standard classifier behavior plus modern zero-shot capabilities: probabilities over constrained choices, no autoregressive generation, no invalid classes, inference-time labels — things zero-shot/NLI classifiers, embedding models, cross-encoders and rerankers have done for years.
Two core criticisms: Jev's impressive comparisons are against LLMs, when of course a specialized classifier beats autoregressive generation on speed and cost — the meaningful comparison is against strong classifier baselines. The ICLR BTZSC benchmark covers dozens of zero-shot classifiers across 22 datasets, yet Jev hasn't been properly benchmarked in that landscape.
Where community comparisons exist, the story is far less magical: on Banking77, BGE-small + logistic regression reached 93.3% accuracy at 9ms locally versus 83.2% for Jev (jev-baselines-eval on GitHub). The author concludes public evidence only shows what was already known — specialized classifiers are cheaper and faster than LLMs for classification. Whether it's a new paradigm hinges on their unpublished architecture and RLCD training method.
More from Companies & People
- Debate: Should AI Safety Researchers Hire Personal Assistants? — herbiebradley · 2026-09-24
- Betting on Unpolished Talent Is Riskier but Full of Diamonds — anshulkundaje · 2026-09-24
- Education's Twin Missions: Backing Proven Talent and Unpolished Diamonds — anshulkundaje · 2026-09-24
- Finding Latent Talent Without Opportunity Is Education's Other Mission — anshulkundaje · 2026-09-24
- Three years ago Claude was a punchline: Garry Tan amplifies lesson on iterating to PMF — garrytan · 2026-09-24
- Anthropic's Claude-led gene editing discovery disputed for ignoring prior work — soumitrashukla9 · 2026-09-24