Advanced Reasoning Models Quietly Quit: RELIC Benchmark Exposes LLM Flaws

tallinzen · x · 2026-08-05

Researchers introduced RELIC, a new evaluation framework that tests LLMs' complex reasoning by asking them to determine if a string belongs to a context-free language (CFL) presented in-context.

Experiments reveal that even the most advanced reasoning models perform poorly on RELIC. As task complexity increases, models fail to scale their inference compute appropriately and actually reduce the number of reasoning tokens they use. This decrease in compute is accompanied by a shift in reasoning strategy, where models move from identifying and implementing algorithmic solutions to simply guessing. For models whose full completions go uninspected, this manifests as "quiet quitting."

Related event: Researchers Debate LLM Reasoning Limits as RELIC Framework Exposes Failures(6 posts)→

Original post →

More from Research

Research channel →