HCOMP paper: fluent LLM explanations drive over-reliance in hate speech moderation
windx0303 · x · 2026-09-30
The Bing Tuo team traveled from Australia to present their Honorable Mention paper 'Easy to Read, Easy to Trust' at HCOMP 2026. The study examines how processing fluency in LLM explanations drives over-reliance in hate speech moderation, finding that easier-to-read explanations tend to increase uncritical trust in model outputs — a caution for human-AI collaboration in moderation workflows.
More from Safety
- Noted hacker Ben Hawkes joins Anthropic to lead Frontier Red Team's cybersecurity mission — logangraham · 2026-09-30
- Grok and Gemini power America .gov, the US government's new AI answer site — elonmusk · 2026-09-30
- Palisade AI Seeks Frontier Lab Employees for Interview Project, Welcomes Risk Skeptics — davidmanheim · 2026-09-30
- XBOW's AI agent finds and exploits a Linux kernel 0-day, achieving local root privilege escalation — moyix · 2026-09-30
- Early AI incidents are evidence of bad security and alignment practices, researcher argues — jessi_cata · 2026-09-30
- New paper: mechanism design could have prevented the Hugging Face incident — joshgans · 2026-09-30