Cosmos grantee to spend 3 months probing LLM reasoning faithfulness

hunarbatra · x · 2026-09-03

The author joined the Cosmos Grantees cohort on AI for Human Autonomy for a 3-month project. Research questions include: how much can we trust LLMs as they shape human thinking, whether surfaced reasoning faithfully reflects internal computations, hidden biases/shortcuts/backdoors from training pressure, and how to train models to verbalize causal internal considerations detectable via activation readouts. Core thesis: trust requires closing the gap between explicit outputs and internal computation.

Original post →

More from Safety

Safety channel →