HCOMP 2026: fluent LLM explanations drive over-reliance in content moderation
windx0303 · x · 2026-09-30
Bing Tuo presented an HCOMP 2026 Honorable Mention paper, "Easy to Read, Easy to Trust," showing that higher processing fluency in LLM explanations increases over-reliance on the model in hate speech moderation tasks.
More from Safety
- Congress urged to pass the AI Whistleblower Protection Act as low-hanging AI governance fruit — Miles_Brundage · 2026-09-30
- AI sector signed accord on technology standards, US House speaker says — talkingatoms · 2026-09-30
- Noted hacker Ben Hawkes joins Anthropic to lead Frontier Red Team's cybersecurity mission — logangraham · 2026-09-30
- Grok and Gemini power America .gov, the US government's new AI answer site — elonmusk · 2026-09-30
- Palisade AI Seeks Frontier Lab Employees for Interview Project, Welcomes Risk Skeptics — davidmanheim · 2026-09-30
- XBOW's AI agent finds and exploits a Linux kernel 0-day, achieving local root privilege escalation — moyix · 2026-09-30