Mapping the Cognitive Structure of LLMs
Zhongxiang Sun · hf · 2026-07-14
This paper introduces NeuroCogMap, an attempt to organize the internal features of Large Language Models through the lens of cognitive neuroscience, mapping these features to interpretable functions, cognitive abilities, and hierarchical structures.
Key findings include:
- Models contain relatively stable, semantically consistent functional regions that are partially reproducible across different models.
- Failure modes such as hallucinations, bias, refusal failures, and sycophancy correspond to dysfunctions in distinct representational or behavioral control systems.
- These internal signatures can be leveraged for mechanism-driven detection and targeted interventions.
- The framework also improves predictions of cortical responses during human natural language processing, particularly in higher-order association cortices.
- The authors further claim these internal signatures can reflect and correct latent strategies in classical human decision-making models.
Overall, this presents a system-level framework connecting "internal model structures—behavioral failures—human cognition."
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22