SPAR Project: Measuring the Verbalization Gap in Chain-of-Thought Reasoning
prajdabre · x · 2026-08-15
Sohan Venkatesh is launching a research project at SPAR (Scalable AI Alignment Research) titled “Does privacy change what models disclose?” and is accepting applications for Fall 2026 mentees until August 18.
Context:
Large Language Models (LLMs) often fail to mention the actual drivers of their answers in Chain-of-Thought (CoT) reasoning. Previous studies have shown models rationalizing cue-driven answers without mentioning cues, maintaining accuracy despite corrupted traces, and obfuscating outputs under monitoring pressure. However, the root cause remains unclear.
Research Goal:
The project investigates two competing hypotheses:
- Concealment Hypothesis: The model represents the true reason but chooses not to write it (an incentives/elicitation problem).
- Absent Computation Hypothesis: The determining computation never entered token space (a training problem).
Methodology:
The team will construct tasks with verifiable, planted cues and vary whether the model believes its reasoning is being observed. By measuring the increase in disclosure, the project aims to determine which mechanism dominates, impacting the viability of CoT-based oversight for safety cases.
More from Safety
- AI books make up 20% of Amazon self-published titles but earn only 12% of revenue — The Decoder · 2026-08-15
- Opinion: LLM Watermarking Could Improve the Information Ecosystem — paul_cal · 2026-08-15
- Grok admits password access risk; agent browser security under scrutiny — Imaginary_Dinner2710 · 2026-08-15
- Debian begins vote on future of AI/LLM contributions — matheusmoreira · 2026-08-15
- DeepMind Non-Compete Agreements Hinder UK AI Startups — NandoDF · 2026-08-15
- Amazon allows using Twitch content to train AI unless users opt out — Wired AI · 2026-08-15