New BPJ jailbreak bypasses top defenses with single-bit black-box attacks
StephenLCasper · x · 2026-08-13
A new paper introduces Boundary Point Jailbreaking (BPJ), a fully automated black-box attack that uses only a single bit of information per query—whether the classifier flags the interaction—to evade the strongest industry safeguards, including Constitutional Classifiers and GPT-5's input classifier. BPJ converts a target harmful string into a curriculum of intermediate targets and actively selects boundary points to detect small improvements, achieving the first automated attack against GPT-5's classifier without human seeds.
More from Safety
- DeepMind Policy Lead and Experts Launch AI Governance Publication — round · 2026-08-13
- Anthropic Report Finds Current Retraining Programs Insufficient for AI Job Displacement — paulnovosad · 2026-08-13
- Smuggling 'Ignore Previous Instructions' with Invisible Characters: New Prompt Injection Trick — GiiTZzz · 2026-08-13
- Paper proposes safety case framework for AI misuse safeguards — StephenLCasper · 2026-08-13
- Google deploys production-ready probes for Gemini, tackling long-context shifts — StephenLCasper · 2026-08-13
- Anthropic's constitutional classifiers withstand 1,700 hours of red-teaming with 0.5% error — StephenLCasper · 2026-08-13