PLSemanticsBench Shows LLMs Struggle to Reason From Rules
PLSemanticsBench tests whether models can reason from the formal semantics of a brand-new programming language. While some models reach up to 90% accuracy, performance drops by 40 to 60 points when familiar symbols like “+” are remapped to new meanings, suggesting they rely more on pretrained lexical associations than on the provided rules.
2026-07-07 ~ 2026-07-07 · 4 related posts
- PLSemanticsBench:大模型能读懂新语言的语义规则吗? — jessyjli · 2026-07-07
- PLSemanticsBench:把➕重定义为减法,性能暴跌40-60分 — jessyjli · 2026-07-07
- PLSemanticsBench:大模型靠词汇联想而非给定规则推理 — jessyjli · 2026-07-07
- PLSemanticsBench:熟悉符号被翻转含义比陌生符号更伤性能 — jessyjli · 2026-07-07