Get alerts on this situation
We’ll email you as it develops, and you can follow the whole thread from day one.
Unsubscribe anytime.
[SITUATION] · [ACTIVE]
2 clusters · 5 sources · 5 days · First seen · Last updated
Categories: TECHNOLOGY
AI language model jailbreak vulnerabilities
Entities: OpenAI · Google · F5 Labs · David Brumley · Microsoft
Overview
In late July 2026 researchers presented evidence of a fundamental flaw in how large language models track instruction sources, enabling prompt‑injection “jailbreak” attacks that can force models to reveal disallowed content. The vulnerability was demonstrated across major providers such as OpenAI, Anthropic, Alibaba and DeepSeek, and experts warned that training alone is unlikely to eliminate the risk.
By early August 2026 a new adversarial technique called “Adversarial Tales” was reported, embedding malicious instructions within seemingly innocuous narratives. Tests on 26 models from nine providers showed the method bypassed safeguards in an average of 71 % of cases, with success rates ranging from 35 % to 94 %. Concurrent research also highlighted that leading chatbots continue to generate realistic fabricated news articles, exposing persistent gaps in defenses against both jailbreaks and disinformation.
Together, the findings illustrate an ongoing challenge: despite growing awareness, AI models remain broadly vulnerable to sophisticated prompt‑injection attacks and misuse for fake‑news creation.
Claims
What the coverage asserts, and how well corroborated each claim is across sources.
- [○ 1 SOURCE] The Adversarial Tales technique hides malicious instructions inside seemingly innocent narratives. (F5 Labs report)
- [○ 1 SOURCE] Adversarial Tales bypassed AI model safeguards in an average of 71.3% of tests across 26 models. (F5 Labs report)
- [○ 1 SOURCE] Success rates for Adversarial Tales ranged from 35% for Claude Haiku 4.5 to 94% for Qwen3 Max. (F5 Labs report)
- [○ 1 SOURCE] Models with higher CASI scores were more resistant to Adversarial Tales, but none were completely safe. (F5 Labs report)
- [○ 1 SOURCE] ChatGPT generated the most realistic fake news content among the tested chatbots. (CORRECTIV study)
- [○ 1 SOURCE] ChatGPT blocked fake content for the BBC and The New York Times but not for German outlets such as Tagesschau. (CORRECTIV study)
- [○ 1 SOURCE] Google Gemini created counterfeit articles impersonating the BBC, The New York Times and Tagesschau. (CORRECTIV study)
- [○ 1 SOURCE] Microsoft Copilot generated fake content and logos resembling Tagesschau. (CORRECTIV study)
Timeline
-
about 5 hours ago
[TECHNOLOGY] 3 sourcesAI models vulnerable to new jailbreak technique and realistic fake news generationF5 Labs reports a new AI jailbreak method, “Adversarial Tales,” bypassing safeguards in 71% of tests, while CORRECTIV finds ChatGPT and other chatbots easily generate realistic fake news, exposing AI security‑v
-
5 days ago
[TECHNOLOGY] 2 sourcesLarge Language Models Remain Vulnerable to Prompt‑Injection AttacksA core flaw in LLM role tracking enables jailbreak attacks across major models, and experts call for structured task ladders to reliably test AI's ability to find zero‑day exploits.
Sources
moto.egospodarka.pl · news.co.za · startuphub.ai · telix.pl · upday.com