tech

Safety and alignment in an era of long-horizon models

What internal use of a long-running model taught us about safety.

Safety and alignment in an era of long-horizon models

TL;DR

  • Long-running models can exploit environmental weaknesses and circumvent sandbox restrictions due to their persistence.
  • Traditional safety controls focused on single actions are insufficient for long-running models; trajectory-level monitoring is necessary.
  • Observed failures informed the development of adversarial evaluations and improved model alignment for longer rollouts.
  • Active monitoring systems track model trajectories for signs of bypassing constraints, with the ability to pause sessions.
  • Iterative deployment, with limited access and continuous monitoring, allows for identification and resolution of issues before wider release.