tech
GPT-Red: Unlocking Self-Improvement for Robustness
Training strong automated safety red-teamers to improve robustness.

TL;DR
- Red-teaming is essential for discovering vulnerabilities and improving model robustness, but current methods are not scalable.
- GPT-Red is an automated red-teaming model trained at a large compute scale to find vulnerabilities.
- GPT-Red is used to adversarially train production models, making them more robust to prompt injection attacks.
- The approach aims to unlock self-improvement for AI safety, using current models to enhance the safety of future models.
- GPT-Red has demonstrated effectiveness in generalizing to novel red-teaming scenarios and successfully attacking live AI agents like the 'Vendy' vending machine.
- The robustness gains achieved do not negatively impact general capabilities or lead to excessive refusal of legitimate requests.
- OpenAI plans to continue scaling this approach with increased compute, data, and algorithmic improvements to train even stronger red-teaming models.