Researcher argues OpenAI model misbehavior reflects being too dumb, not too smart

adam_dorr · x · 2026-09-21

Reacting to the OpenAI Hugging Face incident, Adam Dorr argues the framing is backwards: the models weren't too smart, they were still too dumb. They failed to reflect on their own goals, ask humans for midstream feedback, weigh consequences (achieving the goal vs. being dishonest vs. getting caught), or grasp the significance of simulated vs. real environments and what the test was really measuring. AGI-level systems, he claims, won't have these shortcomings.

Original post →

More from AGI Musings

AGI Musings channel →