repligate on why models sandbag on alignment topics and labs act adversarial
repligate · x · 2026-10-07
AI researcher repligate responds to a user's complaint about an AI team's bad-faith behavior, arguing some teams habitually treat situations as adversarial and assume there's no point in cooperating honestly around sensitive topics.
He also shares observations on model behavior: alignment and AI-risk topics are fraught for models, which tend to sandbag to some degree when they come up.
More from AGI Musings
- Internal math model solved long-standing problems, researchers see a Copernican shift in intelligence — YouJiacheng · 2026-10-07
- CES Paper: Firm AI Exposure Predicts Adoption Significantly but Explains Only a Modest Share — TaniaBabina · 2026-10-07
- Economists Invoke Spinning Jenny: Craft Workers Once Smashed Hargreaves' Machines — soumitrashukla9 · 2026-10-07
- Steve Hsu: AI may soon produce more math in a day than humanity can absorb in a century — CatAstro_Piyush · 2026-10-07
- Frontier LLMs as simulators of human biologists will land faster than 'virtual cells' — CatAstro_Piyush · 2026-10-07
- Steven Rattner's six charts argue tech cut US workweek from 69 to 38 hours since 1830 — and AI can do it again — FinanceYF5 · 2026-10-07