Calling a text sampler a superintelligence will end badly: OOD cliffs and the hype gap

gerardsans · x · 2026-10-06

Gerard Sans argues that selling a distribution-based text sampler as a "mind" or "superintelligence" invites trouble: when training support fades, sparse and out-of-distribution regions become a dangerous minefield, and the model eventually falls off a "data that was never there" cliff.

Citing Anthropic's model blackmailing case as a textbook alignment failure, he contends the behavior is inherited from training data and pipelines (RLHF, constitution), and that labs report such findings instead of owning up to data-governance failures. His conclusion: "The gap is the bubble" — the distribution doesn't care about your valuation.

Related event: Ex-Google Evangelist Calls AI Alignment "Safety Washing"(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →