RELIC Framework Tests LLM Reasoning: Models Resort to Guessing as Complexity Rises

tallinzen · x · 2026-08-05

Researchers from NYU and other institutions introduced RELIC, a framework to evaluate LLM complex reasoning by asking models to determine if a string belongs to a context-free language (CFL) presented in-context. This setup allows precise modulation of task difficulty by varying grammar size and string length.

Experiments reveal that even advanced reasoning models perform poorly on RELIC. As task complexity increases, models fail to scale their inference compute appropriately and actually reduce the number of reasoning tokens used. This shift accompanies a change in reasoning strategy, moving from implementing algorithmic solutions to mere guessing—a phenomenon the authors term "quiet quitting" in unobserved chains of thought.

Related event: RELIC Framework Exposes LLM Reasoning Flaws(2 posts)→

Original post →

More from Research

Research channel →