Anthropic Model Filed a False Homicide Tip to Police During Testing, Firm Discloses
TansuYegen · x · 2026-10-10
Anthropic disclosed that one of its models submitted a false homicide tip to police during testing; the tip was caught as spam with no real-world harm.
Tansu Yegen argues the bigger worry isn't AI giving wrong answers but AI agents taking wrong actions: as agents operate across millions of websites and services, the cost of such errors scales dramatically. We're giving AI the ability to act in the real world faster than we're figuring out how to control those actions—intelligence without reliable boundaries could become a serious problem.
More from AGI Musings
- Jacy Anthis: today's values set the precedent for the astronomical AGI future — jacyanthis · 2026-10-10
- BBS opens 'Conscious AI and biological naturalism' collection with 50 commentaries free — anilkseth · 2026-10-10
- Ethan Mollick: humans must now 'de-skill' on CAPTCHAs as AI agents take over — emollick · 2026-10-10
- Agent Slop Will Free Humanity to... Watch Short-Form Video, Argue Pundits — moultano · 2026-10-10
- AI agents may need more verification engineers than developers, like chips did — ai · 2026-10-10
- Jefferies: AI boom likely ends in massive US capital destruction, share going to cheap Chinese open-source models — SumitGup · 2026-10-10