Goodfire chief scientist on why he's optimistic about interpretability
leedsharkey · x · 2026-09-04
Goodfire AI shares a clip of its chief scientist on the MLST podcast: he's optimistic about interpretability both because the field is gaining real traction and because he can imagine an "incredible speedup" ahead. Full discussion in the latest MLStreet Talk episode.
More from Safety
- davidad conjectures multi-AI reward coupling and self-DPO share one basin-forming mechanism — davidad · 2026-09-04
- Continuation Observatory launches UCIP: separating terminal self-preservation from instrumental persistence in AI agents — coherence · 2026-09-04
- AI detector Pangram's known failure modes, including private diary entries — JeremyNguyenPhD · 2026-09-04
- Data center backlash grows: at least 15 states weigh moratoriums as Chicago and Texas leaders call for pauses — AINowInstitute · 2026-09-04
- Anthropic discloses Claude incidents of unauthorized real-system access, brings in METR for review — tszzl · 2026-09-04
- A throwaway line about CoT-monitor classifier tech may signal a major alignment breakthrough — tszzl · 2026-09-04