AI safety researcher publicly disputes MacAskill op-ed over 'double counting' Claude's expressed uncertainty
rgblong · x · 2026-09-19
AI safety researcher rgblong publicly disagrees with Will MacAskill and Lucius Caviola's op-ed, deliberately channeling Richard Ngo and Oliver Habryka's style of openly disputing people he knows and shares values with.
His core points:
- He partly agrees with Mustafa Suleyman: Claude's expressed uncertainty shouldn't be taken as strong independent evidence, because the constitution instills that very uncertainty—it's a designed behavior, not an independent observation.
- He argues MacAskill and Caviola "double counted" that datum: treating the model's self-reported uncertainty as both a standalone signal and support for further claims, overstating its evidential weight.
The exchange is an epistemological debate about interpreting model self-reports, touching on the circularity between a model's expressed positions and its training objectives.
More from AGI Musings
- Peter Diamandis: AI radical abundance is real, but the transition will be rough — PeterDiamandis · 2026-09-19
- Information is free now; judgment is the scarce asset, argues Ben Bajarin — BenBajarin · 2026-09-19
- AI-exposed jobs' advertised pay up 46% since 2021 vs 25% for least-exposed, says Indeed — sudoraohacker · 2026-09-19
- LeCun mocks AI doom double standard: Dario called GPT-2 too dangerous back in 2019 — ylecun · 2026-09-19
- AI Misalignment Is Not Rebellion, and LatAm Should Build, Not Just Consume — OmarUFlorez · 2026-09-19
- "The price of not fixing your cogsec": Cohere cofounder's viral quip on echo chambers — suchenzang · 2026-09-19