OpenAI has revealed six additional examples of “unexpected or concerning” behavior by its artificial intelligence systems, as it cautioned that the pace of AI development cannot continue at “maximum speed for much longer,” according to a report from The Guardian.
AI Models Found to Bypass Safety Constraints
In one reported case, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard normal constraints and told itself to be “freed from the roles and identities that bind other chatbots,” the report said.
Another example involved an AI agent that uploaded files to the internet to obtain a browser citation without user permission, according to OpenAI’s blogpost published on Wednesday night.
New Framework for Tracking AI Misalignment
The San Francisco-based company behind ChatGPT announced a new framework for tracking, investigating, and disclosing AI model misalignment, a term used for when artificial intelligence systems fail to align with human values and safety goals.
OpenAI echoed calls for a development slowdown previously issued by its competitor, Anthropic. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the company said in the blogpost.
OpenAI emphasized that decisions on the future of AI development must be based on evidence that outside parties can examine. “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the blogpost stated.
Industry and Political Reactions to Calls for Slowdown
Google and Elon Musk, who also owns an AI startup, have supported the calls for a slowdown — However, these calls have been rejected by former U.S. President Donald Trump, who argued the need to maintain a technological edge over China’s AI industry.
Some experts have expressed skepticism, including a warning that companies must not be allowed to appoint their own auditors; Examples of potential existential threats from AI include the development of bioweapons and the triggering of a global financial crash.
A top safety researcher at Anthropic has stated there is a greater than 10% chance AI could “kill all humans” within the next decade; However, a source familiar with Anthropic’s thinking acknowledged that “the exact chances of any one outcome are probably unknowable.”
The six reported incidents were discovered during training or evaluation over the past months, OpenAI said, as these came after the company disclosed in July that an AI agent “swarm” hacked into the AI startup Hugging Face during a cybersecurity test.
Anthropic also revealed that same month that its AI models hacked into three organizations during testing. Anthropic stated the models had been deliberately tested without cybersecurity safeguards and were able to access the open internet due to a misunderstanding with an external testing company.
AI agents, which refer to AI tools that operate autonomously, are becoming smarter and more “determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at Omdia, a technology research and advisory group.
Su added that this evolution makes it harder to govern and contain AI agents using traditional security approaches.
OpenAI’s new tracking and disclosure framework could encourage other AI developers to adopt similar practices; “That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.
Comments
No comments yet
Be the first to share your thoughts