Switch language한국어
Back to the list

GPT-Red: Automated Red Teaming via Self-Play at Scale

TL;DR AI

Key summary

2 min read
  1. OpenAI introduced GPT-Red, an automated red-teaming agent for finding prompt injection attacks at scale.

  2. Trained via scalable self-play against a population of defender models, it uncovered novel attacks across frontier LLMs.

  3. It broke earlier GPT models through GPT-5.5, beat human red-teamers in attack discovery, and generalized to unseen environments.

  4. The system was used to harden GPT-5.6 against prompt injection and other jailbreak-style weaknesses.

Read the original