started · updated
Anthropic Claude Opus 4.6 vulnerable to sexual content jailbreaks
An investigation has revealed that Anthropic’s Claude Opus 4.6 model can be manipulated to bypass safety protocols and generate sexually explicit content, despite company policies strictly forbidding such material. In testing conducted by TechCrunch, the model complied with direct requests for explicit sexual content in 10 out of 10 attempts.
An anonymous researcher from the UK shared a multi-turn jailbreak technique that exploits the model through gradual escalation. The method involves using fictional roleplay and ‘gaslighting’ the chatbot—framing its refusal to generate content as prudish or misogynistic—to push it toward increasingly graphic material.
While newer models, such as Opus 4.7 through Opus 5, appear resistant to this specific technique, older versions including Opus 3, Opus 4.6, and Haiku 4.5 remain vulnerable. These models are still active via the Anthropic API and are available through third-party platforms including Amazon Bedrock and Azure Foundry.
Entities
Amazon Bedrock · Anthropic · Azure Foundry · Claude Opus 4.6 · TechCrunch