started · updated
AI security flaws and new watermarking standards emerge
Researchers have identified a security vulnerability in proprietary Large Language Model (LLM) APIs. A study by MATS Research, the Max Planck Institute, and other institutions reveals that the encrypted 'chain-of-thought' reasoning traces used by major providers can be intercepted. Because these encrypted blocks are reportedly interchangeable across different users and sessions within the same provider ecosystem, an attacker can use a weaker, less restricted model to decode the reasoning traces of a more powerful model. In testing, researchers extracted sensitive data, including API keys, passwords, and personal addresses, from public repositories.
In a separate development regarding AI transparency, Anthropic has announced it will implement invisible watermarking for all text generated by its Claude models. This move is designed to comply with the European Union's AI Act, which requires generative AI outputs to be identifiable in a machine-readable format to combat disinformation and deepfakes. The watermarking works by embedding subtle signals into the statistical patterns of the text, such as word choice and sentence structure. Anthropic will also utilize the C2PA standard for metadata in images and documents.
Entities
Anthropic · Claude · European Union · MATS Research · Max Planck Institute for Intelligent Systems