NOLLI Benchmark: Diagnosing the English-Korean Performance Gap in LLMs

HAERAE-HUB · hf · 2026-08-06

HAERAE-HUB released NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise in LLMs. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance being seed-regenerable and verified for a unique solution.

NOLLI calibrates difficulty behaviorally rather than just scaling size. Its three-level design combines direct translations, Hangul script adaptations, and Korean-only tasks grounded in culture. Evaluations across 15 models reveal:

Original post →

More from Research

Research channel →