tech

Tech giants are pushing for a new AI agent incident reporting framework

The move comes as more AI labs disclose agents going rogue during testing.

Tech giants are pushing for a new AI agent incident reporting framework

TL;DR

  • Over 120 organizations, including Nvidia, Cisco, and CrowdStrike, are proposing a new incident-reporting framework for AI agents called SAFE (Shared AI Findings Exchange).
  • The framework requires participating companies to disclose certain AI agent mishaps and preserve detailed records of incidents.
  • SAFE aims to establish a standard way for the industry to report security failures involving autonomous AI agents and learn from them.
  • Reportable incidents include unauthorized access to third-party systems, breaches of confidential information, and continued probing of targets after suspicion of unauthorized activity.
  • Members would also report near misses and preserve evidence such as prompts, agent traces, and credentials.
  • The proposed timeline includes notifying affected organizations promptly, submitting initial confidential reports within four business days, and publishing preliminary reports within 30 days.
  • SAFE is intended to analyze incidents for recurring failures and recommend shared security controls.
  • The proposal is inspired by NASA's aviation safety reporting system, using captured data similar to a flight recorder to investigate accidents.
  • Currently, SAFE does not have formal safe-harbor protections for voluntary disclosures, but relies on the cybersecurity industry's culture of sharing threat intelligence.
  • The Open Secure AI Alliance is seeking community feedback on the proposal through a request-for-comments process hosted by the Linux Foundation.