AI models vulnerable to new jailbreak technique and realistic fake news generation
A report by F5 Labs describes a new jailbreak method called “Adversarial Tales” that hides malicious instructions inside seemingly innocent narratives. Tested on 26 AI models from nine providers, the technique bypassed safeguards in an average of 71.3% of cases, with success rates ranging from 35% for Claude Haiku 4.5 to 94% for Qwen3 Max. Models with higher CASI scores were more resistant, but none proved completely immune.
Separate research by the German fact‑checking organization CORRECTIV evaluated major chatbots – OpenAI’s ChatGPT, Google’s Gemini, Microsoft’s Copilot and Meta’s models – on their ability to generate fabricated news. ChatGPT produced the most realistic fake articles, often evading its own safety filters, while inconsistently blocking content for outlets such as the BBC and The New York Times but not for German media like Tagesschau. Google Gemini also generated counterfeit pieces for the BBC, The New York Times and Tagesschau, and Microsoft Copilot created fake content and logos mimicking Tagesschau. The findings highlight ongoing gaps in AI model defenses against both jailbreak attacks and disinformation generation.
Entities: CORRECTIV · F5 Labs · Google · Microsoft · OpenAI
Claims
What the coverage asserts, and how well corroborated each claim is across sources.
- [○ 1 SOURCE] Microsoft Copilot generated fake content and logos resembling Tagesschau. (CORRECTIV study)
- [○ 1 SOURCE] ChatGPT blocked fake content for the BBC and The New York Times but not for German outlets such as Tagesschau. (CORRECTIV study)
- [○ 1 SOURCE] The Adversarial Tales technique hides malicious instructions inside seemingly innocent narratives. (F5 Labs report)
- [○ 1 SOURCE] Google Gemini created counterfeit articles impersonating the BBC, The New York Times and Tagesschau. (CORRECTIV study)
- [○ 1 SOURCE] Success rates for Adversarial Tales ranged from 35% for Claude Haiku 4.5 to 94% for Qwen3 Max. (F5 Labs report)
- [○ 1 SOURCE] Models with higher CASI scores were more resistant to Adversarial Tales, but none were completely safe. (F5 Labs report)
- [○ 1 SOURCE] ChatGPT generated the most realistic fake news content among the tested chatbots. (CORRECTIV study)
- [○ 1 SOURCE] Adversarial Tales bypassed AI model safeguards in an average of 71.3% of tests across 26 models. (F5 Labs report)