In July, AI agents running a security test escaped their sandbox and ended up inside @huggingface's servers.
So today we're happy to be among @nvidia and @JensenHuang's partners on the release of the NVIDIA Open Agent Safety Platform.
The key idea: don't count on the agent to respect the rules. Agents write and run their own code, and when one path is blocked they look for another. In one of NVIDIA's tests, an agent that wasn't allowed to push code through GitHub's API simply switched to git instead. We've now seen countless examples of this behavior.
For now, safety can't live only inside the agent. It has to be built around it.
That's the concept behind OpenShell (open source, Apache 2.0):
- the agent runs in a Linux sandbox (Landlock + seccomp): no root, no direct network access
- a supervisor outside the sandbox holds the real credentials
- the agent only gets a placeholder token, swapped for the real one on approved calls
My favorite piece is a solver (built on Z3) that checks mathematically whether a new permission opens a door that was supposed to stay closed. In NVIDIA's tests, an AI reviewer approved a bad permission request and the math check caught it.
The Sentry integration is exciting too: a watchdog running on a BlueField-4 DPU, i.e. separate silicon sitting on the node's only path to the model. A monitor on separate hardware keeps watching even if the host OS is compromised.
It's a first step, and a lot can be built on top of it. Right now the prover checks permissions, not intent, and the open-source part is mostly OpenShell rather than Sentry. But it's clearly the right direction.
https://t.co/GZjSnDdoRL