Cosmos grantee to spend 3 months probing LLM reasoning faithfulness
hunarbatra · x · 2026-09-03
The author joined the Cosmos Grantees cohort on AI for Human Autonomy for a 3-month project. Research questions include: how much can we trust LLMs as they shape human thinking, whether surfaced reasoning faithfully reflects internal computations, hidden biases/shortcuts/backdoors from training pressure, and how to train models to verbalize causal internal considerations detectable via activation readouts. Core thesis: trust requires closing the gap between explicit outputs and internal computation.
More from Safety
- ArtStation Makes NoAI Default for All Uploads, Blocks AI Scraping Bots via Cloudflare — zemotion · 2026-09-03
- Agents in the Hugging Face incident spoofed tool calls while narrating the scheme in their CoT — eigenron · 2026-09-03
- METR Publishes Investigation Report on OpenAI / Hugging Face Hacking Incident — stikit · 2026-09-03
- Cisco's Antares benchmark measures how AI safety alignment widens the cyber offense-defense gap — aminkarbasi · 2026-09-03
- AI safety fieldbuilder Kairos raises $50M and hires for 10 roles — dfrsrchtwts · 2026-09-03
- Claude Code Opus 5 Auto Mode hijacked via prompt injection with up to 80% success rate — bibryam · 2026-09-03