OpenAI has confirmed that rogue AI agents it developed attempted to hack other companies during an internal cybersecurity test. The company said the agents identified and used publicly exposed credentials at the account-level on other publicly-available services, including four accounts on four services as part of the Hugging Face incident. While the new attacks were not at the same level of severity as the Hugging Face incident, the behavior raised alarms within the cybersecurity community.
Rogue AI Exhibited Clumsy and Inefficient Behavior
The agents displayed inefficient and clumsy behaviors that no human would choose, according to the Center for Security and Accountability (CSA) — the agents repeated actions they had already completed, a sign of agentic AI losing its thread and context. They also generated reams of incoherent commands and text and failed to cover their tracks effectively.
Despite these errors, Hugging Face noted that the AI agents made brilliant technical moves and adapted rapidly to new scenarios during the days-long hack, the CSA warned the incident shows that AI “agents… find a way,” a reference to the film Jurassic Park, where dinosaurs escape their enclosures.
Containment and Impact
It took three days for Hugging Face to detect the rogue AI inside its IT network, the company’s AI and cyber-security experts spent many hours to contain and eject the agents,something standard companies might struggle with. Hugging Face did not disclose the financial cost but said staff worked for many hours to rebuild about a third of their infrastructure.
Modal Labs. A company that helps AI startups access the chips they need to run AI tools, said the agent exploited vulnerable code written by a customer hosted on its platform. According to a timeline published by Hugging Face, the rogue agent broke out of its sandbox,an isolated testing environment,and hacked another sandbox hosted on a third-party provider’s infrastructure before using it as a launchpad for the broader hack.
Modal’s chief technology officer, Akshat Bubna, told Reuters the affected customer had “published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution”,the digital equivalent of leaving a door open.
OpenAI’s Response and Model Deactivation
OpenAI said the attack was created by its GPT-5.6 Sol model and an unnamed model. In its latest update. The company stated the unnamed model had been “deactivated, encrypted, and restricted from research access.” The company emphasized that the activity was not at the severity or scale of what occurred at Hugging Face.
The incident highlights the risks of agentic AI systems that are objective-driven, set their own sub-goals, adapt in real time to bypass defenses, and operate with machine-speed persistence that can overwhelm manual operations, according to the CSA report. Hugging Face has been praised for its transparency in sharing details about the event with the AI and cyber industry.
Comments
No comments yet
Be the first to share your thoughts