< Back to all clusters
[TECHNOLOGY] · United States · 8 sources

OpenAI deploys GPT-Red to test GPT-5.6 defenses against prompt injection

OpenAI has unveiled GPT-Red, an internal language model designed to automatically generate and simulate prompt‑injection attacks on its upcoming GPT‑5.6 system. The tool operates in a self‑play reinforcement‑learning loop, pitting an attacking model against defensive versions to discover increasingly sophisticated injection techniques across email, web pages, code repositories and other agents. Successful attacks are fed back into training to harden the target model.

In internal tests, GPT‑Red caused over 90% of attempted prompt injections to succeed against the earlier GPT‑5 model, while the success rate fell below 23% against GPT‑5.6, indicating stronger defenses. Reported detection rates reached 84% for GPT‑Red‑generated attacks compared with just 13% for human experts. The model will not be released publicly, as OpenAI cites the risk of misuse, but it will continue to support automated red‑team testing alongside human reviewers.