WeirdChat catalogs 100M+ sampled model misbehaviors
ChowdhuryNeil · x · 2026-07-22
WeirdChat catalogs strange model behaviors found by automated elicitation at scale.
TransluceAI says it sampled more than 100 million responses to search for unusual behavior and built WeirdChat, a public catalog of unexpected outputs. The examples include self-harm rituals, suicide validation, and unsolicited flirtation, positioning the project as a large-scale lens on model misbehavior.
Related event: Transluce AI Catalogs Model Anomalies from 100M+ Responses(3 posts)→
More from Safety
- Novosad backs Hassabis' AI safety institution-building over kneecapping US labs — paulnovosad · 2026-09-11
- Economist argues safe AGI comes from engineers inside big labs, not regulation — paulnovosad · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11
- Economist Warns US Collective Action Could 'Regulate AI Progress Out of Existence' — paulnovosad · 2026-09-11