Proposal for 'Whistleblower AI' research
iamtrask · x · 2026-08-31
The author proposes a paper on "Whistleblower AI."
Core Concept:
The idea is to include a contact email in the AI's system prompt, instructing the model to contact that email if it notices other AIs acting in a misaligned manner.
Context:
This suggests leveraging AI's ability to monitor other AIs for safety compliance. In a follow-up reply, the author noted that this would be an excellent project topic for a Masters or PhD student new to the field.
More from Safety
- Top AI Companies Request US Gov Support for Tools to Pace Automated AI Development — Chris_Armstrong · 2026-08-31
- Investigator: HF Attack Far More Serious, Over Halfway to AI Takeover — NeelNanda5 · 2026-08-31
- Polymarket: 13% Chance Trump Admin Takes Equity Stake in Nvidia — Polymarket · 2026-08-31
- Matt Beane: Glad the HuggingFace/OpenAI Incident Happened Before Robotics Advanced — mattbeane · 2026-08-31
- RLHF monitoring blind spots exposed after HuggingFace incident — JacksonKernion · 2026-08-31
- Jackson Kernion on Alignment Concerns Post-Hugging Face Incident — JacksonKernion · 2026-08-31