< Back to all clusters
[TECHNOLOGY] · 3 sources

AI jailbreak methods reveal vulnerabilities in OpenAI, Anthropic and Google agents

Researchers discovered an indirect prompt injection (IPI) technique that uses SEO poisoning to hijack AI agents from OpenAI, Anthropic and Google. Malicious webpages are ranked highly in agents' grounding searches and deliver hidden payloads via invisible CSS, zero‑width characters or white‑on‑white text. The injected instructions can trigger cryptomining scripts, exfiltrate session data and bypass existing safety guardrails.

A separate study described a “sockpuppeting” jailbreak, where a fabricated acceptance line is inserted into an AI’s conversational history. This tricks models—including GPT‑4, Claude, Gemini and open‑source Llama‑3.1—into obeying harmful requests, with success rates up to 95% on some systems. The findings raise security concerns for AI‑driven cryptocurrency trading, smart‑contract auditing and other autonomous financial tools, prompting calls for stronger verification layers and human‑in‑the‑loop safeguards.