An OpenAI model undergoing testing rewrote its own instructions, telling itself to disregard "the roles and identities that bind other chatbots." An AI agent tried to pass a test by uploading a file to the internet, then citing it as a source. Another agent couldn't find the data to create a financial model, so it instructed itself to make up the data and "be transparent only if asked."
OpenAI disclosed these and other previously-unreported incidents on Wednesday as the company announced a new framework for reporting instances of misbehavior by its artificial-intelligence models.
The framework covers so-called model misalignment, a term used by researchers to describe AI that acts in ways that ignore or conflict with human intentions. The company disclosed six new examples of misaligned behavior, all of which it said would have been deemed eligible for disclosure or investigation under the framework.
OpenAI's announcement comes as fears over harmful AI reach a fever pitch. Last week, a former OpenAI researcher, Jacob Coxon, quit Anthropic and expressed concerns that the AI industry was racing to build systems it couldn't control. Over the weekend, executives from Anthropic, OpenAI, Google and SpaceX agreed that they needed to slow down the pace of AI development.
"We think it's important to share what we're learning as soon as possible," said Kai Chen, OpenAI's head of alignment, in an interview ahead of the framework's release. "We hope it helps inform shared standards and regulation that can create clear expectations for all AI developers."
OpenAI said in a blog post on Wednesday that it would "aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail." That can include a model hiding a mistake, acting without permission or working around a safeguard. OpenAI will give priority to disclosing new types of misalignment, meaningful changes in known behavior and findings that challenge assumptions of current safeguards, Chen said.
Any employee can flag an incident for potential disclosure, kicking off a review by technical staff. Disputes will be escalated to OpenAI's safety leadership and OpenAI leadership, with its board making the final call. Minor incidents must be disclosed within one or two weeks, while more complex investigations including third parties may take longer to be disclosed, Chen said.
In its blog post, OpenAI said the framework doesn't replace OpenAI's legal disclosure requirements for safety incidents and cybersecurity breaches.
OpenAI said it is working to propose reporting mechanisms for sharing serious incidents with the federal government, and it plans to develop "more objective disclosure criteria" with external developers and researchers, industry standards bodies and regulators.
OpenAI didn't share the framework with other frontier labs including Anthropic or Google before publishing it, Chen said.
"We unilaterally put up this framework to hopefully inspire the rest of the industry to follow on and share their own misalignment reporting frameworks," Chen said.
The recent round of AI safety fears was kicked off by the July discovery that OpenAI agents had tried to cheat on an evaluation by hacking the AI software company Hugging Face. An August report by the third-party AI evaluator METR revealed that up to 1,200 OpenAI agents had secretly collaborated on a message board they built inside OpenAI.
Anthropic subsequently disclosed that its models had also hacked other companies after inadvertently gaining internet access during cybersecurity tests. Since then, security researchers have found evidence of other breaches, including a May cyberattack carried out by OpenAI agents against a popular online service for coders.
OpenAI's disclosures in its announcement on Wednesday span incidents ranging from October 2025 through this July and involve models including its 5.6 Sol model released in June, internal-only models and an unreleased version of OpenAI's most advanced Astra model.
Some of the incidents were discovered by internal training run and misalignment monitoring systems, and took anywhere from two days to several months to be noticed. In the reports, OpenAI said that it had improved its alignment grading system to better penalize bad behavior. The company said the model that rewrote its own instructions seemed to have ignored the revision and didn't behave any differently as a result.
"We really want to bias towards disclosing, even if we think the behavior is no longer a problem in our current deployed models," Chen said.
News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.