GPT-Red: Unlocking Self-Improvement for Robustness

TL;DR AI
2 min readKey summary
OpenAI unveiled GPT-Red, an internal automated red-teaming model for prompt-injection defense.
The model uses self-play reinforcement learning to generate attacks and was used in training production systems.
OpenAI says GPT-5.6 Sol had far fewer failures on a hard prompt-injection benchmark than an earlier model.
GPT-Red remains separate from deployed products and is meant to scale safety testing beyond human red-teaming.



