Using Committee Prompting for Content Moderation: LLMs Stuck in Infinite Loops

pbloemesquire · x · 2026-08-06

A blog post explores AI alignment issues when using committee prompting for content moderation.

The article presents a case where a committee of AI agents is tasked with judging whether a controversial tweet violates rules. The agents get stuck in a loop:

While simulating a committee is a common alternative to direct prompting, it can lead to unproductive overthinking and stagnation when dealing with complex moderation criteria.

Original post →

More from Safety

Safety channel →