Claude Code autoresearch loop discovers jailbreaks beating 30+ GCG attacks, accepted at NeurIPS 2026

maksym_andr · x · 2026-09-25

Researchers deployed Claude Code in an autoresearch loop to autonomously discover novel jailbreaking algorithms. The resulting attacks beat 30+ existing GCG-style attacks, aided by AutoML hyperparameter tuning.

The work, dubbed Claudini, has been accepted at NeurIPS 2026 — cited as one of the first successful applications of autoresearch to (hill-climbable) safety tasks, a strong sign that incremental safety and security research can now be automated.

Original post →

More from Safety

Safety channel →