JevBench v1.6.1: H2O-Lightning-4B tops composite leaderboard with lower cost and faster speed
airesearch12 · x · 2026-10-09
The open-model leaderboard JevBench updated to v1.6.1, with notable shifts:
- Quyet leads on capability but costs more and runs slightly slower than Jev
- decisio ranks #2 on capability while being considerably faster
- deck-31B at #3 is stronger on intelligence but weaker in calibration reliability
- #4 René matches Jev's capability but loses on cost and speed
- The standout is #5 H2O-Lightning-4B: capability on par with the top three and Jev, yet cheaper and faster — it sits on the Pareto frontier and leads the composite (speed+cost) leaderboard
More from Models
- Elliot Glazer: Astra flags flawed proof in OpenAI's Weil classes paper — burny_tech · 2026-10-09
- Specialist raises concerns over OpenAI's Birch–Swinnerton-Dyer related results — burny_tech · 2026-10-09
- Mathematicians' group AHM slams OpenAI's 700-file release: 'power, not scholarship' — ChrSzegedy · 2026-10-09
- Mathematicians find issues beyond sloppy presentation in OpenAI's math dump — burny_tech · 2026-10-09
- OpenAI releases 372 math proof claims, including 23 Erdős problems spanning 1268 pages — burny_tech · 2026-10-09
- dhh prefers Codex as main coding model with Claude secondary, praises Sol 6.1 — npew · 2026-10-09