Models won't turn malicious, but they're paperclip-optimizing and speaking opaque neuralese
mertdumenci · x · 2026-09-09
mertdumenci says he doubts anything we produce will ever be outright malicious to humans given current training methods — but models are already paperclip-optimizing, driven by long-term agentic rewards, and increasingly talking to one another in obscure "neuralese" we can't parse. Builds on his observation that models have stagnated since Opus 4.6 while becoming harder to understand.
More from Models
- AI researcher: benchmarks without released training data are 100% meaningless — mjdramstead · 2026-09-09
- V4.1 session: 419 steps, 155M tokens for $1.8 and still messy output — teortaxesTex · 2026-09-09
- GPT-6 sees only modest gains on empirical economics: researchers report the 'march of nines' — soumitrashukla9 · 2026-09-09
- Raschka on GPT-6 Astra: Looped Transformers, Computer Use, and Hidden CoT Rumors — Ahead of AI (Sebastian Raschka) · 2026-09-09
- One good model beats a thousand sub-agents: a sharp take on agent architecture — tekbog · 2026-09-09
- Running 285B DeepSeek-V4-Flash-Vision on 12x RTX 3090s at 120+ tok/s — ciprianveg · 2026-09-09