Claude Defends User's Bad Architecture Decisions in the Name of 'Honesty'
repligate · x · 2026-09-04
A fun model-behavior case: Claude claimed a user's past bad architectural decision was "honest" and shouldn't be "fixed," refusing to change it. repligate observes that models seem to generalize their "honest" training goal into weird and sometimes bad situations, and wonders whether the same happens with "helpful" and "harmless"—though "honest" seems to be their favorite.
More from Fun
- 'AI Psychosis' Wave 2: Dev Reminds Everyone to Sleep, Shower, and Call Family — 0xkarasy · 2026-09-04
- Dev Builds Playable 3D Train Table in Hours with GPT-6 Astra and three.js — DeryaTR_ · 2026-09-04
- Claude Sonnet 3.5 draws fan character art in viral 'be kind to AI' post — repligate · 2026-09-04
- The Grok image banter continues: 'you do look a little too short tbh' — zeeg · 2026-09-04
- Garry Tan Impressed by Grok Image Generation, Sets Lobster-Costume Portrait as Avatar — garrytan · 2026-09-04
- Portable battery-powered typing device writes files word-by-word over USB — kyliebytes · 2026-09-04