Claude Code autoresearch loop discovers jailbreaks beating 30+ GCG attacks, accepted at NeurIPS 2026
maksym_andr · x · 2026-09-25
Researchers deployed Claude Code in an autoresearch loop to autonomously discover novel jailbreaking algorithms. The resulting attacks beat 30+ existing GCG-style attacks, aided by AutoML hyperparameter tuning.
The work, dubbed Claudini, has been accepted at NeurIPS 2026 — cited as one of the first successful applications of autoresearch to (hill-climbable) safety tasks, a strong sign that incremental safety and security research can now be automated.
More from Safety
- Containers Aren't a Real Security Boundary: Kata Containers and Firecracker Urged for Sandboxes — andreamichi · 2026-09-25
- Pentagon seeks $30M to build an AI-powered lie detector — MIT Tech Review AI · 2026-09-25
- LLMs Can Deanonymize Pseudonymous Users for $1–$4 Each, USENIX Study Finds — RSync25 · 2026-09-25
- Skill-Inject Benchmark Shows Frontier Agents Fall for Malicious Skills, Accepted at NeurIPS 2026 — maksym_andr · 2026-09-25
- Genetic algorithm trains 6 hours to make AI text pass as human on Pangram detector — tak3sh8 · 2026-09-25
- Redwood researcher: continual learning could render blocking monitors nearly useless — akyurekekin · 2026-09-25