NOLLI Benchmark: Diagnosing the English-Korean Performance Gap in LLMs
HAERAE-HUB · hf · 2026-08-06
HAERAE-HUB released NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise in LLMs. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance being seed-regenerable and verified for a unique solution.
NOLLI calibrates difficulty behaviorally rather than just scaling size. Its three-level design combines direct translations, Hangul script adaptations, and Korean-only tasks grounded in culture. Evaluations across 15 models reveal:
- Matched English-Korean accuracy is statistically equivalent within a ±10 pp margin, suggesting little cost from presentation language alone.
- Writing-system-intensive tasks show sharp gaps: Korean Cipher falls behind English by up to 68.7 pp.
- A salient size measure fails to predict empirical difficulty in 7 of 15 types, making model size an unreliable proxy.
More from Research
- UC Berkeley Introduces RHI: Optimizing Agent Harnesses to Cut Inference Costs by 60% — ceciletamura · 2026-08-06
- LLMs as Autonomous Cyber Defenders: Multi-Agent Security Research — xuanalogue · 2026-08-06
- Nature Publishes Landmark HCMI: 665 Cancer Organoids from 2,780 Patients Released — anshulkundaje · 2026-08-06
- SKILL-KD: Contrastive Skill Distillation for Weaker LLM Agents — ZhejiangUniversity · 2026-08-06
- Princeton Introduces Skill Entropy to Measure and Boost LLM Cross-Skill Reasoning — princetonu · 2026-08-06
- Tencent Study: VLM Agents Face Severe Safety Risks from Stale Spatial Memory — tencent · 2026-08-06