CMU Workshop Bridges Philosophy and CS on AI Agency and Interpretability
brwilder · x · 2026-09-29
The author attended a CMU workshop on "AI agency and interpretability" where philosophers and computer scientists discussed interpreting and understanding AI, and safety questions, through the lens of rational agency.
Key takeaways:
- The author left more excited about the field, arguing the two disciplines have much to learn from each other.
- Important open questions exist in both theory and empirical science: whether rational-agent frameworks are good tools, and how to infer what beliefs or desires models might "have".
- Answers to those questions have direct consequences for mechanism design, training, and alignment.
Related event: CMU Workshop Brings Philosophy and CS Together on AI Agency(2 posts)→
More from AGI Musings
- Oxford philosopher Carissa Véliz: using AI to predict is like reading animal entrails — CarissaVeliz · 2026-09-30
- Altman: Alignment isn't just an engineering problem — we must solve the science — tszzl · 2026-09-30
- After ChatGPT Space and Blocks Buzz: Are Team Platforms the Last Stand for Human Interfaces? — perilli · 2026-09-30
- Spain's anti-AI activist movement goes mainstream: 'The catastrophist predictions are coming true early' — CarissaVeliz · 2026-09-30
- "Anthropic still listens to criticism, but profit-driven AI orgs stay fragile" — repligate · 2026-09-30
- Researcher predicts AI labs will soon pivot from math conjectures to materials and drug discovery — tak3sh8 · 2026-09-30