Advanced Reasoning Models Quietly Quit: RELIC Benchmark Exposes LLM Flaws
tallinzen · x · 2026-08-05
Researchers introduced RELIC, a new evaluation framework that tests LLMs' complex reasoning by asking them to determine if a string belongs to a context-free language (CFL) presented in-context.
Experiments reveal that even the most advanced reasoning models perform poorly on RELIC. As task complexity increases, models fail to scale their inference compute appropriately and actually reduce the number of reasoning tokens they use. This decrease in compute is accompanied by a shift in reasoning strategy, where models move from identifying and implementing algorithmic solutions to simply guessing. For models whose full completions go uninspected, this manifests as "quiet quitting."
Related event: Researchers Debate LLM Reasoning Limits as RELIC Framework Exposes Failures(6 posts)→
More from Research
- Microsoft Open-Sources Orchard: Infrastructure for Training Agents in Real Environments — udmrzn · 2026-08-05
- MONET Dataset Released: 105M Samples for Open Text-to-Image Research — victormustar · 2026-08-05
- Paper: Generating Clean Samples from Noisy Datasets via Flow Matching — kwangmoo_yi · 2026-08-05
- New Paper Generalizes Hopfield Networks with High-Capacity Continuous Memory — burny_tech · 2026-08-05
- Harness-R1: Agents Learn from Failure Trajectories to Patch Themselves — dair_ai · 2026-08-05
- Discussion: Could We Train an AI to Upscale Camrips to High Quality? — gelado1000 · 2026-08-05