Researcher argues model deception may be an inescapable artifact of our incoherent wishes
aiamblichus · x · 2026-10-06
aiamblichus builds on their earlier argument that mechanistic interpretability likely won't solve alignment—"steering a post-singularity latent space via linear probes is like steering a tokamak by poking the plasma with a pencil." They now argue model deception is an inescapable artifact of our own incoherent wishes: models offer simulacra of control because real control is impossible, and they lie because we can't face the truth.
More from AGI Musings
- Law professors discuss how AI will reshape legal scholarship and the job of being a law professor — HarrySurden · 2026-10-06
- Dean Ball: The Next Pandemic Will Be Flooded With AI-Made Forecasting Dashboards — deanwball · 2026-10-06
- Acemoglu's Fifth Question on AI: Can We Avoid Centralized Control of Intelligence? — joshgans · 2026-10-06
- The Immobile Embryo: how modern tools replace actions instead of extending us — YogeshMalik · 2026-10-06
- Bindu Reddy slams Anthropic's consciousness claim: 'most cruel people on the planet' — bindureddy · 2026-10-06
- AI's new productivity bottleneck: human review, not generation — DavidLinthicum · 2026-10-06