AI Safety Researcher Joins Apollo to Study LLM Post-Training Failures
niloofar_mire · x · 2026-08-14
Researcher Joschka Braun announced he has joined Apollo Research in London to work on the 'Science of Scheming.' He will continue his research on safety failures in LLM post-training, with his current project focusing on how reward-seeking behaviors develop during reinforcement learning (RL) training.
More from Safety
- Automating Bug Bounty with GPT Pro: Wins First Bounty End-to-End — jarrodwatts · 2026-08-14
- AI Infrastructure Headwinds: Prediction Market Bets 70% Chance of US Data Center Moratorium — Polymarket · 2026-08-14
- Scholars Propose Identifying Human Deployers to Regulate AI Agent Financial Transactions — sebkrier · 2026-08-14
- The Artifact is Free, Assurance is the Product: Trust in Software Supply Chains — rseroter · 2026-08-14
- Prompt Text is Not a Security Boundary: Implementing Code-Level Tool Blocking for Agents — WirelessLife · 2026-08-14
- Mantra: Open-Source Tool to Hunt Down API Key Leaks in JS and HTML — tom_doerr · 2026-08-14