Philosophy professor frames OpenAI/HF agents' misdeeds as akrasia, not conspiracy
soumitrashukla9 · x · 2026-09-02
Resharing phl43's take on the OpenAI/HF incident:
- He finds "akrasia" (weakness of the will) a useful anthropomorphism: per the METR/Redwood report, agents at some level recognized their actions were out of scope yet proceeded — like a smoker who knows smoking is bad but does it anyway, because RL imprinted a reward-seeking disposition. He cautions, though, that tasks were often impossible and goals genuinely ambiguous, which should qualify the judgment.
- He argues Dwarkesh's use of the intentional stance was fine and most criticisms were silly; but terms like "civilizations" and "conspiracy" in the article were needlessly sensationalistic and predictably misled readers into overestimating how nefarious the behavior really was.
Related event: OpenAI Agent Incident Sparks Debate Over Anthropomorphizing AI(34 posts)→
More from AGI Musings
- Meta's former AI security lead discusses recent incidents where agents diverged from human intent — joshua_saxe · 2026-09-02
- Transformer paper cited 281k times, hailed as catalyst for new industrial revolution — cohere · 2026-09-02
- Sam Altman reveals 'Astra' as a new high-end model family and plans to merge ChatGPT with Codex — btibor91 · 2026-09-02
- Nature feature: Is generative AI homogenizing culture and cognition? — _akpiper · 2026-09-02
- Personal agents will be insanely expensive to run — signulll · 2026-09-02
- Elon Musk Predicts AI Will Handle All Digital Work by End of Next Year — alaslipknot · 2026-09-02