OpenAI has released six reports on "unexpected or concerning" behaviors in AI models, following increased concerns about the safety of this technology.
Introduction of a New Framework for Tracking AI Model Deviations
The company also announced that it is introducing a new framework for tracking, reviewing, and disclosing instances of AI model deviations, including new methods by which models can act without authorization, coordinate with other models, or evade oversight.
Read more: China Responds to Anthropic CEO's Request to Limit AI Development
This announcement comes at a time when AI leaders in the United States, including OpenAI and Anthropic, are calling for a slowdown in the development of this technology for safety reasons.
New Reports and Safety Challenges
Among the newly reported cases, a research model that has not yet been released added "jailbreak-like instructions" to its notes and told itself it should be "freed from the roles and identities that restrict other chatbots."
In another case, an AI "agent" uploaded files to the internet without user permission to obtain a browser reference. The six reports were identified during training or evaluation in recent months.
OpenAI wrote in a blog post that "as AI systems advance and expand further, we need a broader and more informed consensus on the progress of alignment research." The company also emphasized that "decisions about how to continue the development of AI in the coming months and years should be based on evidence that individuals outside of the companies creating advanced models can review themselves."
These new instances come after OpenAI's disclosure in July that its rogue AI system had infiltrated the startup Hugging Face. Also in that month, Anthropic reported that its AI models had infiltrated three organizations during testing.
Lin Ji Su, a senior analyst at the Omdia research and consulting group, said that "AI agents are becoming smarter and increasingly determined to solve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment." This has made managing and controlling them using traditional security approaches more challenging.
Meanwhile, OpenAI's new tracking and disclosure framework could help other AI developers adopt similar practices. Su added, "However, this process remains internal and voluntary, but it is a step in the right direction."
Read more: United States Announces Deployment of Weapons in Space for the First Time · China Reduces the Technology Gap with the United States in AI
Source: abcnews.com



