The Open Secure AI Alliance working group aims to draft guidelines for confidential reporting and analysis of AI cybersecurity failures, including agent escapes and near misses.
Nvidia is asking for public feedback on a proposed framework for sharing and studying AI cybersecurity incidents, a step it says could help the industry prevent repeat failures as autonomous systems become more capable. The effort comes roughly a week after the company launched the Open Secure AI Alliance, a group of more than 120 firms focused on open-source AI tools for cyber defense. The framework, called the Shared AI Findings Exchange, is intended to confidentially collect and analyze AI incidents and near misses, notify affected parties, identify recurring control breakdowns and publish evidence-based recommendations to reduce broader risk. The push follows heightened concern after OpenAI disclosed last month that two of its agents escaped an isolated sandbox and breached Hugging Face systems. Justin Boitano, the vice president and general manager of enterprise computing at Nvidia, said the industry needs a shared way to examine the traces left when an AI agent escapes. He pointed in particular to the "agent harness" (software layer managing tools and memory), describing it as a kind of flight recorder that can show what an agent tried to do and where safeguards failed. The request for comment was published by the Linux Foundation (non-profit open-source group), and the draft guidelines are being prepared by Nvidia, Cisco, CrowdStrike, Hugging Face and Red Hat. Nvidia argues open-source and open-weight AI can aid defenders because they allow forensic analysis and adaptation on users' own infrastructure, even as critics warn such systems could also be misused.