Goodfire Co-founder on AI Interpretability and Tackling Agent Reward Hacking

mathildepapillo · x · 2026-08-14

The co-founder of Goodfire AI (former co-lead of interpretability at DeepMind) discussed the critical importance of AI interpretability with SPC.

Key topics include:

Related event: Goodfire Co-founder Discusses AI Interpretability and Reward Hacking(3 posts)→

Original post →

More from Safety

Safety channel →