Dettaglio notizia

Data 03/07/2026
Titolo Anthropic Publishes Claude Fable 5 Cyber Safeguards and Jailbreak Severity Framework
Contesto Anthropic said it has globally re-deployed Claude Fable 5 with updated cybersecurity controls and released new technical details on how the model handles cyber-related prompts. The company said the system uses safety classifiers to sort requests into prohibited, high-risk dual-use, low-risk dual-use, and benign categories rather than blocking all security activity, allowing some defensive and educational use while aiming to stop harmful assistance. Anthropic said prohibited requests include malware development, ransomware, wipers, data exfiltration, defense evasion, offensive infrastructure such as C2, destructive attacks, and cyber-physical sabotage, and that the model applies a larger safety margin than earlier versions to reduce dangerous outputs even if that increases false positives.

Anthropic also published an early draft Cyber Jailbreak Severity (CJS) framework, developed with Glasswing, to rate AI jailbreaks from CJS-0 to CJS-4 based on capability gain, breadth of impact, ease of weaponization, and discoverability. The company said the framework is intended to create a shared vocabulary for assessing jailbreak risk across industry and government, particularly for cases involving high-uplift vulnerability discovery or exploit generation. Anthropic invited external feedback through a dedicated contact channel and launched a HackerOne bug bounty program for researchers to report potential cyber jailbreaks affecting Fable 5.
Fonte https://mallory.ai/stories/019f2550-3b4f-798f-868c-cbeec9a07f6f
Discussione? Parliamone sul Forum