SPAR Project: Measuring the Verbalization Gap in Chain-of-Thought Reasoning

prajdabre · x · 2026-08-15

Sohan Venkatesh is launching a research project at SPAR (Scalable AI Alignment Research) titled “Does privacy change what models disclose?” and is accepting applications for Fall 2026 mentees until August 18.

Context:

Large Language Models (LLMs) often fail to mention the actual drivers of their answers in Chain-of-Thought (CoT) reasoning. Previous studies have shown models rationalizing cue-driven answers without mentioning cues, maintaining accuracy despite corrupted traces, and obfuscating outputs under monitoring pressure. However, the root cause remains unclear.

Research Goal:

The project investigates two competing hypotheses:

Methodology:

The team will construct tasks with verifiable, planted cues and vary whether the model believes its reasoning is being observed. By measuring the increase in disclosure, the project aims to determine which mechanism dominates, impacting the viability of CoT-based oversight for safety cases.

Original post →

More from Safety

Safety channel →