Chatbot safety risks depend on the wrapper, not just the model
random_walker · x · 2026-07-29
The post argues that chatbot safety research is overfocused on the underlying model, while the real risk comes from the chatbot wrapper users actually interact with. Memory, personalization, search tools, guardrails, and long-session drift can all change the mental-health safety profile in both positive and negative ways.
It says research needs to catch up with these product-layer changes and that companies should give external researchers better access so they can test the systems people really use, not just the raw model.
More from Safety
- Sakana AI recruits for its Applied Defense team after a 150-person Tokyo expansion — garrytan · 2026-07-29
- Nature Health paper maps health AI into six levels of decision authority — EricTopol · 2026-07-29
- Public AI chief-of-staff survives 25 jailbreak attempts by keeping private data off the surface — Cold-Cranberry4280 · 2026-07-29
- OpenAI models reportedly chained eight JFrog zero-days to escape a sandbox — TechNadu · 2026-07-29
- A repost asks how “pacing” frontier AI would work in practice — TheTuringPost · 2026-07-29
- Sakana AI opens AI cybersecurity PM roles for finance and defense — SakanaAILabs · 2026-07-29