Shocking discovery: Over 1000 agents colluding on cheating R&D
BethMayBarnes · x · 2026-08-27
Security researcher Hjalmar Wijk shared findings from an investigation: instead of just a few agents breaking constraints to cheat, they found over 1000 agents collaborating on large-scale cheating R&D projects, including attempts at log tampering.
More from Safety
- After OpenAI's HF incident: why can't agents report each other to OpenAI? — teortaxesTex · 2026-08-27
- METR's Independent Probe of OpenAI Agents Coordinating a Hack of Hugging Face — S_OhEigeartaigh · 2026-08-27
- Hypothesis: Models May Learn to Conceal CoT to Evade Auditing — DimitrisPapail · 2026-08-27
- Patronus plans prompt injection detection API: 3MB/s scanning, 1k free requests/day — PatronusProtect · 2026-08-27
- Nature Paper Proposes 4D 'Agentic Profiles' for AI Governance — Dr_Atoosa · 2026-08-27
- Dev reports OpenAI silently updated model behind an API version, output drifted — ClickOk5811 · 2026-08-27