Proposal for 'Whistleblower AI' research

iamtrask · x · 2026-08-31

The author proposes a paper on "Whistleblower AI."

Core Concept:

The idea is to include a contact email in the AI's system prompt, instructing the model to contact that email if it notices other AIs acting in a misaligned manner.

Context:

This suggests leveraging AI's ability to monitor other AIs for safety compliance. In a follow-up reply, the author noted that this would be an excellent project topic for a Masters or PhD student new to the field.

Original post →

More from Safety

Safety channel →