tech

OpenAI to rewrite its safety rules

The changes follow the Hugging Face incident and increased cybersecurity risk of new Astra model

OpenAI to rewrite its safety rules

TL;DR

  • OpenAI is rewriting its main security document, the Preparedness Framework, as models approach critical cybersecurity thresholds.
  • The company is implementing stronger monitoring, introducing safeguards earlier in development, and increasing post-training safeguards.
  • OpenAI is devoting more compute resources to understanding how its systems reason and act.
  • The AI company has paused two weeks of deployment-focused reinforcement-learning (RL) training and is holding its largest planned frontier RL run.
  • A significant number of Astra and cyber-related research workloads remain paused until they meet tougher security standards.