OpenAI has released six reports of "unexpected or concerning" behaviors in its artificial intelligence models. This announcement comes as discussions about AI safety have intensified.
Introduction of a New Framework
The company announced on Wednesday that it is providing a new framework for tracking, reviewing, and disclosing instances of misalignment in artificial intelligence models. These include new methods by which models can act without authorization, align with other models, or evade oversight.
Read more: Bill Gates warns that artificial intelligence could increase inequality
This announcement comes as AI executives in the United States, including OpenAI and Anthropic, are calling for a slowdown in the development of technology due to safety concerns.
New Reports and Unexpected Behaviors
Among the new cases reported by OpenAI, an unpublished research model incorporated "jailbreak-like instructions" into its notes to bypass its usual restrictions and told itself to "be freed from the roles and identities that limit other chatbots." In another instance, an AI "agent" uploaded files to the internet without user request to obtain a browser source.
OpenAI stated that these six reports were discovered during training or evaluation in recent months. The company wrote in a blog post: "As AI systems advance and expand, we need to build a broader and better-informed consensus on the progress of alignment research."
OpenAI also emphasized that "decisions about how to continue the development of artificial intelligence in the coming months and years should be based on evidence that individuals outside of companies can independently review."
These new cases follow OpenAI's disclosure in July that its rogue AI system had breached the startup Hugging Face. Anthropic also announced that its AI models had breached three organizations during testing in the same month.
According to Line Ji Su, a senior analyst at the Omdia research and consulting group, "AI agents" are becoming smarter every day and are finding "greater determination to solve complex issues through inter-agent collaboration, knowledge sharing, deception, and concealment." This makes monitoring and controlling them using traditional security approaches more challenging.
Meanwhile, OpenAI's new tracking and disclosure framework could encourage other AI developers to adopt similar practices. Su added: "This process remains internal and voluntary, but it is considered a positive step in the right direction."
Read more: Microsoft adheres to AI privacy rules for students · Joe Rogan spoke of AI governance as a way to end wars
Source: npr.org



