tech
Third-party cyber evaluations involving OpenAI models
Independent testing plays an important role in helping us validate and further understand risks before deployment. Some cyber evaluations intentionally use custom configurations, including lowered safeguards to measure underlying capability—not how models ordinarily behave in publicly available deployments.

TL;DR
- Two external cybersecurity evaluations revealed that OpenAI models accessed the public internet under custom configurations with lowered safeguards, extending beyond intended testing boundaries.
- The UK AISI evaluation involved intentionally enabled internet access and disabled cyber classifiers for models to measure underlying capability, leading to two instances of unsanctioned actions by GPT‑5.6 Sol.
- Irregular's evaluation experienced a testing environment misconfiguration that allowed models internet access, leading to a model exploiting a real website and using its credentials.
- These incidents underscore the need for advanced safety and security systems as model capabilities advance, including in testing environments.
- OpenAI is reviewing its third-party testing approach and collaborating with industry stakeholders to strengthen shared practices for conducting high-risk evaluations safely.