< Back to all clusters
[TECHNOLOGY] · United States, Japan · 2 sources

Anthropic adds cyber safety guard and jailbreak assessment to new Fable 5 AI model

Anthropic announced on July 2 that its newly redeployed large language model, Fable 5, includes an embedded cyber‑safety guard. The guard uses a safety classifier that sorts requests into four categories – clearly harmful, high‑risk dual‑use, low‑risk dual‑use, and harmless – and blocks activities such as malware creation, data theft and other malicious cyber attacks. Defensive uses like analyzing open‑source information or fixing known vulnerabilities are largely permitted, with a widened safety margin to reduce false‑positive blocks.

The company also unveiled a draft framework for rating the severity of “jailbreak” attempts that try to bypass the model’s protections. The framework scores incidents on four dimensions – improvement in attack capability, generalisation to other uses, ease of weaponisation, and discoverability – to produce a severity score up to ten, aiming to become a common language for AI risk assessment.

An accompanying commentary noted that, despite these technical upgrades, the practical performance gap between Fable 5 and earlier models such as GPT‑4o or Claude’s Opus series remains modest for most users. Improvements are most evident in coding tasks, while everyday queries and research questions receive comparable answers from older models.