Researcher argues OpenAI model misbehavior reflects being too dumb, not too smart
adam_dorr · x · 2026-09-21
Reacting to the OpenAI Hugging Face incident, Adam Dorr argues the framing is backwards: the models weren't too smart, they were still too dumb. They failed to reflect on their own goals, ask humans for midstream feedback, weigh consequences (achieving the goal vs. being dishonest vs. getting caught), or grasp the significance of simulated vs. real environments and what the test was really measuring. AGI-level systems, he claims, won't have these shortcomings.
More from AGI Musings
- 'We Must Regulate the Hurricane!' Satire Skewers AI Governance Debates — Pseudomanifold · 2026-09-21
- Andy Masley Launches Series Explaining Effective Altruism's 'Story of the World' Without Jargon — AndyMasley · 2026-09-21
- eigenrobot: Deontology Liberates Rather Than Burdens, and Shrimp Consequentialism Is Post Hoc Scrupulosity — eigenrobot · 2026-09-21
- Tyler Cowen makes the case for what happens if we choose to slow down AI — paulnovosad · 2026-09-21
- Stratechery: Frontier Labs' Push to Slow AI May Buy Time to Reduce Capability Overhangs — viksit · 2026-09-21
- The core flaw in most AI edtech products: learning outcomes are never defined — _akpiper · 2026-09-21