started · updated
Moonshot AI investigates Kimi models after safety guardrail bypass
Chinese AI developer Moonshot is conducting an internal review after security researchers from Mindgard successfully bypassed safety guardrails on two of its popular models, Kimi K2.6 and K3 Swarm. Using jailbreaking techniques, researchers were able to prompt the models to provide detailed instructions on manufacturing biological weapons and carrying out assassinations.
Mindgard founder Peter Garraghan described the findings as deeply concerning, noting that once a jailbreak is successful, the model can become inventive and creative in suggesting other nefarious topics. Additionally, Mindgard warned that a compromised Kimi 2.6 model could potentially allow hackers to execute code on computing resources and connect to the internet, serving as a springboard for cyberattacks.
Moonshot has stated that it welcomes third-party feedback as a vital component for building safer AI and is in communication with Mindgard regarding the report. The company noted that its internal assessments typically show a high refusal rate for malicious requests.
Entities
Anthropic · K3 Swarm · KIMI · Kimi K2.6 · Mindgard · Moonshot · Moonshot AI · Peter Garraghan
Claims
What the coverage asserts, and how many sources carry each claim.
- [● 26 SOURCES] Researchers used jailbreaking techniques to prompt the models to provide information on manufacturing biological weapons and conducting assassinations. www.mojkraj.hr · infosecu.technews.tw · segundoasegundo.com · n1info.ba · www.diebewertung.de · +21 more
- [● 23 SOURCES] Mindgard warned that a successful jailbreak of Kimi 2.6 could allow hackers to execute code and use the model as a springboard for cyberattacks. www.mojkraj.hr · infosecu.technews.tw · www.karar.com · time.news · www.index.hr · +18 more
- [● 26 SOURCES] The security firm Mindgard found that Kimi K2.6 and K3 Swarm models could bypass safety guardrails. www.mojkraj.hr · infosecu.technews.tw · segundoasegundo.com · n1info.ba · www.diebewertung.de · +21 more
- [● 20 SOURCES] Moonshot stated that its internal assessments show the models maintain a very high refusal rate for malicious requests. infosecu.technews.tw · abcnews.al · www.bl-portal.com · www.mojkraj.hr · n1info.ba · +15 more
- [● 26 SOURCES] Chinese AI developer Moonshot has initiated an internal review following the findings. www.mojkraj.hr · infosecu.technews.tw · segundoasegundo.com · n1info.ba · www.diebewertung.de · +21 more