NYU's CIA Metric Finds LLMs' Chain-of-Thought Only 44.8-75.9% Aligned With Internal Computation
newyorkuniversity · hf · 2026-10-07
NYU researchers measure and improve the alignment between LLMs' chain-of-thought and their internal computations.
Key contributions
- CIA (CoT-Interpretability Alignment) metric: measures agreement between CoT traces and internal reasoning strategies detected by interpretability tools
- Evaluated on two-hop QA, hint intervention, and integer multiplication across three LLMs, finding limited alignment (44.8-75.9%)
- Post-training with both task accuracy and parametric faithfulness as rewards substantially improves CoT faithfulness while maintaining or improving task accuracy
- Code and data open-sourced: github.com/yihuaihong/CIA-minimal-repro
The work provides both an auditing framework for CoT faithfulness and a pathway toward more trustworthy explicit reasoning.
More from Research
- Converting 10 coding harnesses including Claude Code into RL environments, cutting tool calls 31% — vllm_project · 2026-10-07
- Google DeepMind's IdeaLens reads outlines to detect AI-generated ideas, not AI-written words — rohanpaul_ai · 2026-10-07
- Anima Anandkumar's First UN Visit: The 'Silent Revolution' of AI in Science — AnimaAnandkumar · 2026-10-07
- WRAP: Adversarial Training Tackles Deep Hedging in Nonstationary Markets via DRO — chaumian · 2026-10-07
- Oxford's AI Theorist Autonomously Derives Physical Model for Quantum Material α-RuCl₃ — Oxford · 2026-10-07
- HERMES: Executable Dev-Primitives Boost SWE Agent Performance by 12.4% at Lower Cost — Haibo Jin · 2026-10-07