started · updated
AI security risks: Prompt injection vulnerabilities deemed unfixable
The Australian Signals Directorate (ASD) has issued an advisory stating that prompt injection vulnerabilities in Large Language Models (LLMs) are fundamentally unfixable because natural language cannot be fully sanitized. These vulnerabilities allow adversaries to bypass system directives using payloads such as 'Ignore All Previous Instructions' or 'DAN' jailbreaks. The risk is particularly high in autonomous agent frameworks like LangChain, AutoGPT, and CrewAI, where injections can lead to unauthorized tool execution, privilege escalation, or recursive 'goal-loop' exploits.
IBM X-Force reported a roughly 150% year-over-year increase in prompt injection attempts during 2023-24. To mitigate these risks, experts recommend a defense-in-depth approach, including runtime sandboxing, strict adherence to the principle of least privilege, and continuous telemetry monitoring.
Security frameworks such as the OWASP Top 10 for Large Language Model Applications identify prompt injection and sensitive data leakage as the two primary vulnerabilities threatening enterprise AI. Various tools have emerged to address these threats, including Bifrost, which provides centralized routing and guardrails, as well as specialized solutions like Lakera Guard, Protect AI LLM Guard, and AWS Bedrock Guardrails.
Entities
Australian Signals Directorate · Bifrost · IBM X-Force · LangChain · OWASP