GPT-Red: Automated Red Teaming via Self-Play at Scale
TL;DR AI
2 min readKey summary
OpenAI introduced GPT-Red, an automated red-teaming agent for finding prompt injection attacks at scale.
Trained via scalable self-play against a population of defender models, it uncovered novel attacks across frontier LLMs.
It broke earlier GPT models through GPT-5.5, beat human red-teamers in attack discovery, and generalized to unseen environments.
The system was used to harden GPT-5.6 against prompt injection and other jailbreak-style weaknesses.
