Steve Sinofsky: Strip the anthropomorphism — model misbehavior is just bugs
ylecun · x · 2026-09-17
Steven Sinofsky lays out a framework for reporting model misalignment, echoed by Yann LeCun: strip away anthropomorphic language and behaviors described as "thinking," "cheating," or "ignoring instructions" are simply bugs.
- They may be architectural flaws inherent to LLMs or issues in pre/post-processing; fixes range from trivial to nearly impossible, but the nature is unchanged.
- His analogy: an old SQL report printing leftover buffer content on a NULL result wouldn't be said to "ignore instructions" — it just has a bug.
- Agents misbehaving shouldn't be framed as conscious disobedience; software is just doing dumb stuff it shouldn't.
More from AGI Musings
- Nate Silver: The Sudden Surge in AI Safety Coverage Reflects Long-Undercovered Demand — ShakeelHashim · 2026-09-17
- Paradigm proposes a new future for scientific communication, using the Riemann Hypothesis as an example — tensorqt · 2026-09-17
- NYT covers recursive self-improvement; Schmidhuber's ex-PhD student builds RSI startup Inherent — SchmidhuberAI · 2026-09-17
- Computer vision academia shifts: publish or perish becomes publish, promote, or perish — abursuc · 2026-09-17
- Cloudflare's innovator's dilemma: protecting the human web may cost it the agentic one — EdenEmarco177 · 2026-09-17
- MIT Schwarzman dean Daniel Huttenlocher: humans are to blame for AI failures — AlexTensor · 2026-09-17