LLM Safety Trilemma: Useful Capability, Reliable Safety, and Open Access Cannot Coexist

Pingyu Wu · hf · 2026-08-03

This research highlights a fundamental flaw in current LLM safeguards for dual-use tasks: models decide whether to answer before seeing the actual downstream use, while attackers can easily forge benign requests or interaction histories.

The authors theoretically derive a safety trilemma proving that Useful Capability, Reliable Safety, and Open Access cannot coexist. To break the impasse, the paper proposes integrating trusted credentials that add hard-to-copy information to predict actual downstream usage, establishing a reliable safety floor while preserving openness.

Original post →

More from Safety

Safety channel →