RELIC Framework Tests LLM Reasoning: Models Resort to Guessing as Complexity Rises
tallinzen · x · 2026-08-05
Researchers from NYU and other institutions introduced RELIC, a framework to evaluate LLM complex reasoning by asking models to determine if a string belongs to a context-free language (CFL) presented in-context. This setup allows precise modulation of task difficulty by varying grammar size and string length.
Experiments reveal that even advanced reasoning models perform poorly on RELIC. As task complexity increases, models fail to scale their inference compute appropriately and actually reduce the number of reasoning tokens used. This shift accompanies a change in reasoning strategy, moving from implementing algorithmic solutions to mere guessing—a phenomenon the authors term "quiet quitting" in unobserved chains of thought.
Related event: RELIC Framework Exposes LLM Reasoning Flaws(2 posts)→
More from Research
- Visualization of the Gravner–Griffeath Hexagonal Cellular Automaton — matthen2 · 2026-08-05
- NeurIPS 2026 Calls for Papers on Embodied Spatial Reasoning and World Models — du_yilun · 2026-08-05
- HazyResearch Open-Sources MoK: A Fused MoE Megakernel for NVL72 — HazyResearch · 2026-08-05
- Journal of Economic Perspectives Publishes AI Symposium — TaniaBabina · 2026-08-05
- Silico: Using Motion AI Models to Detect Parkinson's Gait — ninamiolane · 2026-08-05
- Robotics Needs Custom Small VLMs: Fine-tuned Qwen Beats APIs in Cost and Accuracy — ChongZzZhang · 2026-08-05