PLSemanticsBench Shows LLMs Struggle to Reason From Rules

PLSemanticsBench tests whether models can reason from the formal semantics of a brand-new programming language. While some models reach up to 90% accuracy, performance drops by 40 to 60 points when familiar symbols like “+” are remapped to new meanings, suggesting they rely more on pretrained lexical associations than on the provided rules.

2026-07-07 ~ 2026-07-07 · 4 related posts