JevBench Adds 22-Language Breakdown, Small Models Collapse in Japanese and Danish
airesearch12 · x · 2026-10-05
JevBench now scores every model across 22 languages. Early signal: the 31B models hold up (deck-31B scores 97 in German, Quyet 94), while several 4B models drop to chance level in Japanese or Danish.
If your users aren't English-only, check this table before choosing a model.
More from Models
- Why don't modern LLMs know time has passed between messages? — dumierhan · 2026-10-06
- Reflection AI's new text model reportedly pretrained on ~24T tokens, multimodal version expected — nagpalchirag · 2026-10-06
- Early User Reports Anthropic's Opus 5.5 Fills Its Context Window Quickly — rickasaurus · 2026-10-06
- Viral Claude vs GPT Charts Mislead: Claude's "5x Value" Is Mostly Just Higher API Pricing — jdjohnson · 2026-10-06
- Rumor: Zhipu's next open source release GLM 5.5 may beat Claude Opus — bindureddy · 2026-10-06
- Liquid AI's d1 vision decision model matches GPT-6.1 Sol on 4 of 6 tasks at 19x-200x lower cost — JosephJacks_ · 2026-10-06