Anthropic unveils AI safety switch and jailbreak severity framework
Anthropic announced an experimental system called Gradient‑Routed Auxiliary Modules (GRAM) that can compartmentalise and remove dangerous knowledge—such as hacking techniques, virology or nuclear information—from its large language models. The company says GRAM works like a "switch button" that could let future models "forget" high‑risk capabilities while retaining everyday functions, though it has not yet been deployed in production models like Claude.
In a separate move, Anthropic introduced the Cyber Jailbreak Severity (CJS) scale, a four‑tier framework (CJS‑0 to CJS‑4) for rating the risk of AI jailbreaks. The scale evaluates capability gain, breadth, ease of weaponisation and discoverability, and is accompanied by a classification of use cases ranging from prohibited to benign. Anthropic also launched a HackerOne program for researchers to submit jailbreaks for review, aiming to create a common language for AI risk communication with regulators and security teams.