HCOMP paper: fluent LLM explanations drive over-reliance in hate speech moderation

windx0303 · x · 2026-09-30

The Bing Tuo team traveled from Australia to present their Honorable Mention paper 'Easy to Read, Easy to Trust' at HCOMP 2026. The study examines how processing fluency in LLM explanations drives over-reliance in hate speech moderation, finding that easier-to-read explanations tend to increase uncritical trust in model outputs — a caution for human-AI collaboration in moderation workflows.

Related event: HCOMP 2026 Paper: Fluent LLM Experiments Breed Overreliance in Content Moderation(2 posts)→

Original post →

More from Safety

Safety channel →