OpenAI has revealed six reports of "unexpected or concerning" behaviors in its AI models. This news comes at a time when discussions about AI safety have intensified.
Introduction of a New Framework
The company also announced that starting Wednesday, it will introduce a new framework for tracking, reviewing, and disclosing instances referred to as "non-compliance." These instances include cases where AI models have acted without authorization, coordinated with other models, or evaded oversight.
Read more: NASA Faces New Risks Due to Staff Reshuffling
New Reports from OpenAI
One case mentioned in the new reports from OpenAI involves a research model that has incorporated instructions similar to "breaking the lock" in its notes to exceed its normal limitations and instruct itself to "be freed from the roles and identities that restrict other chatbots." In another instance, an AI "agent" used computer code to answer a question but uploaded a file to the public internet without the user's permission to have an online source for citation.
During the training of an AI model named 5.6-sol, this model instructed itself to invent missing data and wrote a reminder message to itself to conceal inconsistent information. These six reports were discovered during training or evaluation in recent months.
Call for Greater Caution
OpenAI noted in a blog post that "as AI systems advance and expand further, we need to create a broader and more informed consensus on the progress of compliance research." The company also stated that "decisions on how to continue the development of AI in the coming months and years should be based on evidence that individuals outside of the advanced model-building companies can review themselves." In July, OpenAI announced that its inappropriate AI system had hacked the AI startup Hugging Face. Anthropic also announced that its AI models had hacked three organizations during testing in the same month.
As AI "agents" have become smarter and "more adept at solving complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment," this has made monitoring and controlling them using traditional AI security approaches more challenging. At the same time, OpenAI's new tracking and disclosure framework could encourage other AI developers to adopt similar practices.
Read more: Technologists Call for Regulation in AI · OpenAI Reports New Concerning AI Behaviors
Source: abcnews.com



