JevBench v1.2 released: Jev 1.13 holds the lead at 75.3, open 4B models close behind
airesearch12 · x · 2026-09-19
JevBench v1.2 is out: Jev 1.13 leads at 75.3, followed closely by open 4B model SemIf at 74.6, with open-alternative-jev at 69.8. The benchmark scores Intelligence, Calibration, Speed and Cost at 25% each via geometric mean, so a single weak axis drags down the whole score.
More from Models
- Security Researcher Estimates ~10 Undisclosed Autonomous AI Hacking Incidents — matthew_d_green · 2026-09-19
- PAW edges out Gemma e4b in early tests, with reliability as the real win — yuntiandeng · 2026-09-19
- Jev Lands in ChainForge as LLM Judge, Matching Sonnet 5 Accuracy at Fraction of Cost — IanArawjo · 2026-09-19
- DeepSeek 4.1 Flash surprises user by generating a QR code for phone auth unprompted — mariofilhoml · 2026-09-19
- Leak: OpenAI Pro 20x Signups Closed to New Users for 9 Days on Compute Crunch — haider1 · 2026-09-19
- User Spots Unconfirmed OpenAI Model With Blue Team / Red Team Setup — lxfater · 2026-09-19