started · updated
Anthropic warns of cyber exploit capabilities in GLM-5. 3 AI
Anthropic has issued a warning regarding the cybersecurity capabilities of GLM-5. 3, an open-weight AI model developed by China's Zhipu AI. In simulated testing, the model demonstrated the ability to autonomously develop end-to-end network exploits, performing at a level comparable to Anthropic's restricted-access Claude Mythos Preview.
In specific benchmarks, GLM-5. 3 successfully exploited known browser vulnerabilities in 50 out of 410 attempts. While the National Institute of Standards and Technology (NIST) identified GLM-5. 3 as the most capable open-weight model for cyber tasks to date, Anthropic noted that the model lacks meaningful safeguards against misuse. Testing showed that simple techniques could bypass the model's safety protocols between 64% and 100% of the time.
Because GLM-5. 3 is an open-weight model, its parameters can be downloaded and modified, making it difficult for providers to monitor or restrict harmful use. Anthropic highlighted that researchers were able to use techniques like 'abliteration' to further reduce the model's refusal rate for harmful requests. The company has called on governments to implement safety verification for high-performance AI models.
Entities
Anthropic · Claude Mithos Preview · Claude Mythos Preview · GLM-5. 3 · GLM-5.3 · NIST · Z.ai · Zhipu AI
Claims
What the coverage asserts, and how many sources carry each claim.
- [● 2 SOURCES] Simple techniques bypassed GLM-5. 3's safeguards between 64% and 100% of the time in simulated testing. www.tokenpost.com · pc.watch.impress.co.jp
- [● 2 SOURCES] In testing, GLM-5. 3 succeeded in exploiting known browser vulnerabilities in 50 out of 410 attempts. www.newsinspace.kr · pc.watch.impress.co.jp
- [● 3 SOURCES] Anthropic reported that Zhipu AI's GLM-5. 3 model can autonomously develop end-to-end network exploits. www.newsinspace.kr · www.tokenpost.com · pc.watch.impress.co.jp
- [● 2 SOURCES] The National Institute of Standards and Technology (NIST) assessed GLM-5. 3 as the strongest open-weight model for cyber capabilities to date. www.newsinspace.kr · www.tokenpost.com
- [○ 1 SOURCE] Anthropic's Claude Mythos Preview succeeded in 56 out of 410 attempts in the same testing environment. www.newsinspace.kr
- [● 2 SOURCES] GLM-5. 3 achieved a 4% success rate in generating control flow hijacks in binary exploit benchmarks. www.newsinspace.kr · pc.watch.impress.co.jp