Jev makes the same errors as flash-tier LLMs, undercutting cascade savings, new study finds
deliprao · x · 2026-09-26
Delip Rao's new arXiv draft examines Jev, a fast, cheap decision model that can drop in for flash-tier LLMs with 1-2 orders of magnitude speed and cost gains.
Key findings:
- Jev makes correlated mistakes with other flash-tier LLMs (Deepseek-V4.1-Flash, GPT-5.6-Luna, Gemini-3.8-Flash).
- Consequence: cascading Jev's "low confidence" outcomes to a larger LLM yields little benefit, since both fail on the same cases.
- Evaluation used 5,000 human-annotated rubric evaluation examples drawn from 9 panels.
- The upside: Jev opens interest in higher-quality open decision models (classifiers).
Takeaway: if the small model and the cascade target share error patterns, the classic cost-saving cascade architecture loses much of its value.
More from Research
- Redwood Research: Astra reasons far better with filler tokens, outside its chain-of-thought — scaling01 · 2026-09-26
- ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS — ysu_nlp · 2026-09-26
- Mathematician Tivadar Danka shares 10 biggest lessons from 20 years in mathematics — TivadarDanka · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Contrastive World Models: latent-space world models without pixel prediction — bonniesjli · 2026-09-26
- Researchers including Google build first complete brain map of a male fruit fly, 166,000+ neurons — burny_tech · 2026-09-26