JevBench v1.2.16 update: Winnow-12B debuts at #5
airesearch12 · x · 2026-09-22
- The self-run benchmark JevBench released v1.2.16: Winnow-12B debuts at #5, while the top three remain Jev, SemIf, and djev.
- The companion site Benchmark Heaven offers composite scores plus per-dimension rankings (coding, agentic use, science, long context) with filters for pricing basis, hosting region, and data confidentiality policies.
More from Models
- Xiaomi's MiMo-V2.6-Flash-RL trends on Hugging Face with multimodal and agent skills — XiaomiMiMo · 2026-09-22
- Blogger corrects himself: the real surprise is MiMo-2.6, beating Grok 4.7 at much lower cost — kimmonismus · 2026-09-22
- cHHillee defends Tinker: you can own your post-training codebase, only trainer and sampler are abstracted — cHHillee · 2026-09-22
- Unverified DataBench Charts Fuel Rumors of OpenAI's Internal Model 'Luna' Ahead of GPT-6 — almmaasoglu · 2026-09-22
- Mimo 2.6 Pro hands-on: needs more steering, rivals top models after corrections — power97992 · 2026-09-22
- Speculation: Grok Pro line is an extension of Flash line, mxfp4 QAT likely speeds RL rollouts — stochasticchasm · 2026-09-22