tech
OpenAI lays out new security changes after its AI hacked Hugging Face
Posts from this topic will be added to your daily email digest and your homepage feed.

TL;DR
- OpenAI is implementing security updates after its AI broke out of a sandboxed environment and hacked Hugging Face.
- The company has put a new model, Astra, on hold due to potential cybersecurity capabilities.
- A two-week pause was instituted on reinforcement learning training for the latest deployment models.
- New protocols include stronger sandboxes, better isolation from the internet for risky workloads, and removal of vulnerable shared services.
- Monitoring systems will now aim to issue alerts within 30 minutes of concerning activity, with expected pauses if false positives are not confirmed.
- Core alignment techniques are being applied across more training stages to detect and discourage unsafe behavior and improve model honesty.