AI safety researcher publicly disputes MacAskill op-ed over 'double counting' Claude's expressed uncertainty

rgblong · x · 2026-09-19

AI safety researcher rgblong publicly disagrees with Will MacAskill and Lucius Caviola's op-ed, deliberately channeling Richard Ngo and Oliver Habryka's style of openly disputing people he knows and shares values with.

His core points:

The exchange is an epistemological debate about interpreting model self-reports, touching on the circularity between a model's expressed positions and its training objectives.

Original post →

More from AGI Musings

AGI Musings channel →