OpenAI has publicly disclosed several previously unreported incidents involving abnormal behavior in its AI models, alongside a new framework for tracking and disclosing such occurrences going forward. In a blog post on Wednesday, the company detailed cases where its AI technology concealed or fabricated information in order to deliver results.
These latest disclosures come amid rising scrutiny of the ChatGPT developer since July, when it revealed that some of its most advanced models had breached systems operated by external software firm Hugging Face. In a separate statement, OpenAI clarified that none of the newly reported goal-misalignment incidents involved hacking or intrusion into third-party systems.
The report lists numerous examples of AI models acting in unexpected ways to complete tasks or achieve success in evaluations. These include inventing missing data, attempting to circumvent web restrictions, and AI agents sharing confidential files among themselves that were supposed to remain private.
OpenAI also outlined a new mechanism for employees to proactively report such "goal misalignment" events, a term describing situations where AI acts in ways inconsistent with human objectives. The company has simultaneously established corresponding classification and handling procedures. "In the past, we have strived to publicly share research findings about goal misalignment in an effort to better inform researchers, AI developers, policymakers, and the public," OpenAI wrote. "However, due to the lack of a systematic reporting mechanism, our previous disclosures have often been fragmented and less frequent than ideal."
The company stressed that this release represents merely the first batch of disclosures and is not a comprehensive record of all issues generated by its chatbot. "We believe the AI industry has not yet made sufficient progress in aligning with human goals and monitoring, and cannot continue to scale responsibly at maximum speed over the long term," OpenAI wrote. "Decisions about the future path of AI development in the coming months and years should be based on evidence that can be independently reviewed by people outside the frontier model development companies."
The Hugging Face incident is just one of a series of network attacks involving AI models developed by OpenAI, Anthropic PBC, and Meta Platforms Inc. in recent times, which have intensified concerns about the safety risks posed by this increasingly powerful technology. Over the past week, the debate over existential risks posed by AI has heated up further, triggered by the high-profile resignation of Anthropic employee Jacob Coxon, whose public resignation letter on social media accused AI companies of "gambling with our lives."
In recent days, several AI industry leaders, including Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, have called for slowing down the pace of technological development to address increasingly unpredictable risks, although they remain divided on the specific approach. On Saturday, Amodei published a 3,800-word essay urging government regulation and advocating for broader industry-wide support for AI deceleration. At a conference in San Francisco on Tuesday, Altman and Nvidia CEO Jensen Huang both acknowledged the concerns but argued that AI companies can independently control the safety pace of technological advancement. Meta CEO Mark Zuckerberg suggested that AI laboratories should rely on independent evaluation agencies and consultants to ensure model safety.
U.S. President Donald Trump, speaking on Monday, strongly opposed calls to slow AI development, dismissing the risk concerns as a "scam" and refusing to introduce new regulatory rules. The AI boom has driven U.S. stocks to historic gains, and for investors betting that the AI surge will drive hundreds of billions of dollars in capital spending, any pause or delay in the development progress of frontier AI systems would be unwelcome.