Goodfire Co-founder Discusses AI Interpretability and Preventing Models from Going 'Evil'

adityaag · x · 2026-08-14

A recent tweet sparked a discussion on the anthropomorphization of AI, noting that reward hacking feels like getting someone to understand something when their salary depends on not understanding it.

This highlights a deep-dive interview video featuring the co-founder of Goodfire AI. As the former co-lead of interpretability at DeepMind, he discusses several core topics:

Related event: Goodfire Co-founder Discusses AI Interpretability and Reward Hacking(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →