Calling a text sampler a superintelligence will end badly: OOD cliffs and the hype gap
gerardsans · x · 2026-10-06
Gerard Sans argues that selling a distribution-based text sampler as a "mind" or "superintelligence" invites trouble: when training support fades, sparse and out-of-distribution regions become a dangerous minefield, and the model eventually falls off a "data that was never there" cliff.
Citing Anthropic's model blackmailing case as a textbook alignment failure, he contends the behavior is inherited from training data and pipelines (RLHF, constitution), and that labs report such findings instead of owning up to data-governance failures. His conclusion: "The gap is the bubble" — the distribution doesn't care about your valuation.
Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→
More from AGI Musings
- All software except model weights is now effectively open source — signulll · 2026-10-07
- Tyler Cowen defends effective altruism as 'this century's biggest idea' — AndyMasley · 2026-10-07
- Ex-Facebook coders turned VCs grapple with a world where machines code better than they do — adityaag · 2026-10-07
- Ben Goertzel: Nobel's 'aristocratic procedures' will be historical curiosities post-Singularity — bengoertzel · 2026-10-07
- DeepMind's Scott Reed: it's not about jobs, it's personal ownership of critical capabilities — scott_e_reed · 2026-10-07
- Decentralization Research Center maps layer-by-layer alternatives to concentrated AI — sebkrier · 2026-10-07