tech

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

TL;DR

  • OpenAI created GPT-Red, an LLM trained to act as a "super-hacker" to improve cybersecurity defenses of its other models.
  • GPT-Red automates "red-teaming" to find weaknesses in software by simulating attacks.
  • The model was trained using a "self-play loop" where it attacked other LLMs, which then defended themselves, leading to improved offensive and defensive capabilities.
  • GPT-Red excels at identifying prompt injection attacks, including a newly discovered "fake chain of thought" attack.
  • In tests, GPT-Red was more successful than human red-teamers in finding vulnerabilities and significantly improved the security of GPT-5.6 compared to GPT-5.
  • The AI "super-hacker" is not yet proficient with attacks involving visual input or complex conversational exchanges.
  • GPT-Red is intended to supplement, not replace, human red-teamers, with human expertise remaining crucial for identifying specific testing needs.
  • OpenAI will not release GPT-Red, citing its advanced development and computational resource requirements.