AI safety paper highlights: reward hacking, RL debate, agent swarms
gasteigerjo · x · 2026-10-07
Johannes Gasteiger's AI Safety at the Frontier newsletter rounds up the best AI safety papers of August & September 2026, covering misaligned reward seekers, reward hacking probes, debate in RL, automated alignment researchers, mind viruses, midtraining tricks, and agent swarms. Full list on his Substack.
More from Safety
- ChatGPT uploads pasted clipboard images before you even hit send, user finds — Innomen · 2026-10-07
- How Abliterated Models Can Get You Pwned — Thrumpwart · 2026-10-07
- OpenAI to watermark ChatGPT outputs by default in the EU under AI Act — Ars Technica AI · 2026-10-07
- Luiza Jarovsky: Beyond a capability threshold, AI alignment is likely impossible — LuizaJarovsky · 2026-10-07
- François Fleuret: IP and regulations are AI's last stand — and they'll be crushed — francoisfleuret · 2026-10-07
- Musubi releases PolicyLM-1.7B, an open-weights model for real-time content moderation — TechCrunch AI · 2026-10-07