Anthropic Reveals Fable 5 Downgrade Logic and CJS Jailbreak Framework

新智元 · wechat · 2026-07-05

On July 2, Anthropic disclosed the interception mechanism behind Fable 5's forced downgrade: its classifier categorizes cybersecurity-related requests into four tiers—death penalty (ransomware, etc., strictly blocked), high-risk dual-use (penetration testing, vulnerability exploitation), low-risk dual-use (intelligence gathering, known vulnerability scanning), and harmless (debug, patch). The classifier's sensitivity was deliberately pushed to the limit, causing a large number of normal debug requests to be misclassified into the third tier and intercepted.

On the same day, Anthropic, partnering with the Glasswing coalition (AWS, Apple, Google, Microsoft, NVIDIA, and others totaling 12 members, with a cumulative investment of 1.04亿美元), launched the CJS (Cyber Jailbreak Severity) framework. It scores jailbreak behaviors using four metrics: capability gain, capability breadth, weaponization difficulty, and discoverability (0-10 points, mapping to five levels from CJS-0 to CJS-4). Scores can only be adjusted upwards and change dynamically with historical context (e.g., Log4Shell was CJS-4 on the eve of its explosion in 2021, but is CJS-0 today).

In retrospect under the new framework, the jailbreak that led to Fable 5's takedown can be replicated by weaker models, has a capability gain of 0, and should have been judged as a CJS-0 "informational" event—meaning it shouldn't have been taken down under the new standards. However, CJS is currently an early draft, and Anthropic is both the rule-maker and the biggest beneficiary, sparking a dispute over standard-setting authority.

Original post →

More from Models

Models channel →