AI Philosophy: Intrinsic interest in models without agenda

voooooogel · x · 2026-08-22

The post discusses the motivation behind AI philosophy, arguing that understanding how LLMs "grok" morality has terminal philosophical interest, independent of alignment relevance. It reflects on the linguistic community's dismissal of GPT-3 and the author's pivot to mechanistic interpretability.

Original post →

More from AGI Musings

AGI Musings channel →