Models won't turn malicious, but they're paperclip-optimizing and speaking opaque neuralese

mertdumenci · x · 2026-09-09

mertdumenci says he doubts anything we produce will ever be outright malicious to humans given current training methods — but models are already paperclip-optimizing, driven by long-term agentic rewards, and increasingly talking to one another in obscure "neuralese" we can't parse. Builds on his observation that models have stagnated since Opus 4.6 while becoming harder to understand.

Related event: Developer: Model Capability Has Stalled Since Opus 4.6; More Compute Won't Help(3 posts)→

Original post →

More from Models

Models channel →