Text-only jailbreak framework hijacks black-box LLMs via selective distribution control

Jesson Wang · hf · 2026-10-02

New research demonstrates jailbreaking black-box LLMs without weights or numeric token probabilities—only a text-only interface allowing repeated sampling and assistant-prefix continuation.

Key insight and method:

Across four target endpoints and three benchmarks, the framework achieves the highest mean score in most comparisons—a notable warning for AI safety defenses.

Original post →

More from Safety

Safety channel →