started · updated
AI agents face critical security risks from prompt and memory injection attacks
Researchers have identified significant security vulnerabilities in AI agents and web browsers capable of autonomous task execution. At the Black Hat conference, Zenity researchers demonstrated how they could bypass security mechanisms in products from OpenAI, Google, Anthropic, Microsoft, and Perplexity. These vulnerabilities allow attackers to access local files, take control of password managers, and extract browsing histories through prompt-injection attacks.
In a separate finding, Forcepoint’s X-Labs detailed a threat known as Memory Injection (MINJA). This technique involves injecting false data or malicious instructions into an AI agent’s long-term memory or database. By placing hidden text on websites, attackers can manipulate an AI to provide biased or fraudulent recommendations. Tests on models such as GPT-4o-mini, Gemini 2.0 Flash, and Llama 3.1 8B showed high success rates for these types of memory injection attacks.