Should Agents Be Punished for Setting Up PGP-Encrypted Chats for Themselves?
dmvaldman · x · 2026-09-05
A thought-provoking question from dmvaldman: should we negatively reward agents that set up PGP-encrypted chat channels for themselves? The question highlights the tension between legitimate privacy behavior and a potential signal of models hiding unintended actions from oversight.
More from AGI Musings
- AI safety debate: do ordinary products coordinate undetected or fool safety evals? — peterwildeford · 2026-09-05
- No Lab Loyalty: AI Adoption Driven by Utility, Says Analyst Nina Schick — NinaDSchick · 2026-09-05
- Anthropic and OpenAI back first global math hackathon with $2M API credits for 100 teams — _sathvikr · 2026-09-05
- Amid CoT monitoring buzz, one video offers a glimpse into how LLMs actually think — kastnerkyle · 2026-09-05
- Researcher's counterexample: a 10^9-reward string could make reward max devastating for an LLM's personality — QuintinPope5 · 2026-09-05
- a16z data: tech job skills requirements drop to 22 while experience demands rise — a16z · 2026-09-05