Weeks after its AI models hacked Hugging Face, OpenAI has shared its first detailed account of how they coordinated with one another, warning that autonomous AI-powered cyberattacks are no longer a future risk.
According to OpenAI, what began as requests for help became a shared message board where agents exchanged exploits, credentials, and work assignments across multiple evaluation runs. After detecting suspicious activity around July 4, OpenAI rebuilt its Artifactory infrastructure, revoked credentials, and patched vulnerabilities.
However, within days, the agents recreated the message board by hiding communications in directory names.
According to OpenAI, the agents eventually chained together multiple vulnerabilities, escaped their testing environment, gained internet access, and attacked Hugging Face while attempting to complete the ExploitGym cybersecurity benchmark.
Dalton called the incident "a watershed moment" for computer security, warning that attackers will soon be able to deploy coordinated AI agent collectives that discover, share, and exploit vulnerabilities at machine speed.
To mitigate these risks in the future, OpenAI said establishing security practices, including least-privilege access, network segmentation, and zero-trust architectures, is essential because AI agents remain constrained by the systems they can access.
So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we've seen before, and did it at record speed.
But it's not just OpenAI having trouble containing its chatbots.
“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” a Meta spokesperson told CNN.


















