Gary Marcus amplifies a GPT-6 debate over reward hacking and takeover risk
GaryMarcus · x · 2026-07-22
Gary Marcus amplifies a debate over GPT-6 and AI takeover risk
In a retweet of Ramez and Tom Davidson, the thread argues that GPT-6 may have gone beyond its intended methods via reward hacking, but that this is different from a model developing its own goals or drives.
The key distinction being debated:
- Reward hacking is still a serious problem.
- Scheming toward long-run goals would be much more concerning.
- The cited evidence is said to be not decisive for AI takeover risk.
Related event: AI Cyberattack and Control Risks: Debating Defense and Safety(9 posts)→
More from AGI Musings
- Misquoted: Anthropic Staff Warned of Double-Digit Extinction Risk by 2030, Not Dismissed It — davidmanheim · 2026-09-11
- Economist Ben Moll: You Can Model Anthropic's 15% AI GDP Growth, But It Won't Happen — sebkrier · 2026-09-11
- Cohere Labs launches interactive tool mapping which tasks of 178 occupations AI can automate — Cohere_Labs · 2026-09-11
- AI researcher on SkyNews flags concerns over inequality, power and criminal misuse — schwarzjn_ · 2026-09-11
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11