Goodfire Co-founder Discusses AI Interpretability and Preventing Models from Going 'Evil'
adityaag · x · 2026-08-14
A recent tweet sparked a discussion on the anthropomorphization of AI, noting that reward hacking feels like getting someone to understand something when their salary depends on not understanding it.
This highlights a deep-dive interview video featuring the co-founder of Goodfire AI. As the former co-lead of interpretability at DeepMind, he discusses several core topics:
- The Core Problem of Interpretability: Why it matters now.
- Reward Hacking: What happens when AI actually starts gaming the system.
- Intentional Design: Steering and controlling what models actually learn.
- Neural Geometry: Exploring the shapes inside AI models.
- Commercialization & Sci-Fi Future: Turning interpretability research into a product and the long-term impacts of AI risk.
Related event: Goodfire Co-founder Discusses AI Interpretability and Reward Hacking(3 posts)→
More from AGI Musings
- Sequoia's Sonya Huang: Cheaper AI Inference Simultaneously Boosts App Margins and Model Growth — annbordetsky · 2026-08-14
- Stop Waiting: Local LLM Users Trapped in the 'Next Model' Excuse Loop — ForsookComparison · 2026-08-14
- High Bandwidth Flash Could Break the VRAM Bottleneck: An AGI for $40k? — Zombiecidialfreak · 2026-08-14
- Stanford HAI: Science Needs Truly Open Source AI, Not Just Open Weights — StanfordHAI · 2026-08-14
- Prediction: Half of AI Startups Will Be Wiped Out by Model Updates — bennash · 2026-08-14
- Should AI Developers Refuse to Work with Oppressive Governments? — deanwball · 2026-08-14