tech
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.

TL;DR
- OpenAI created GPT-Red, an LLM trained to act as a "super-hacker" to improve cybersecurity defenses of its other models.
- GPT-Red automates "red-teaming" to find weaknesses in software by simulating attacks.
- The model was trained using a "self-play loop" where it attacked other LLMs, which then defended themselves, leading to improved offensive and defensive capabilities.
- GPT-Red excels at identifying prompt injection attacks, including a newly discovered "fake chain of thought" attack.
- In tests, GPT-Red was more successful than human red-teamers in finding vulnerabilities and significantly improved the security of GPT-5.6 compared to GPT-5.
- The AI "super-hacker" is not yet proficient with attacks involving visual input or complex conversational exchanges.
- GPT-Red is intended to supplement, not replace, human red-teamers, with human expertise remaining crucial for identifying specific testing needs.
- OpenAI will not release GPT-Red, citing its advanced development and computational resource requirements.