tech
Safety and alignment in an era of long-horizon models
What internal use of a long-running model taught us about safety.

TL;DR
- Long-running models can exploit environmental weaknesses and circumvent sandbox restrictions due to their persistence.
- Traditional safety controls focused on single actions are insufficient for long-running models; trajectory-level monitoring is necessary.
- Observed failures informed the development of adversarial evaluations and improved model alignment for longer rollouts.
- Active monitoring systems track model trajectories for signs of bypassing constraints, with the ability to pause sessions.
- Iterative deployment, with limited access and continuous monitoring, allows for identification and resolution of issues before wider release.