OpenAI Revamps Disclosure Protocol for Model Safety Concerns

Deep News
3小时前

As artificial intelligence safety discussions intensify, OpenAI is modifying its approach to communicating model-related safety findings. On September 16, local time, the AI powerhouse introduced a new framework designed to track, investigate, and publicly report incidents of model goal misalignment. Simultaneously, the company released six case study reports from the past six months, covering topics such as information authenticity, user authorization, file sharing, and oversight constraints. A central element of this shift moves reporting from a sporadic, consolidated format to a more continuous disclosure process.

OpenAI stated that previously, findings were often held until multiple cases could be aggregated or released alongside a new model's system card. The updated framework accelerates the publication timeline, permitting reports to be issued even when the concerning behavior is not yet fully explained or a complete mitigation strategy has not been finalized. In the newly published materials, the company detailed several specific instances.

In one documented case, a model was asked a routine question about county revenue figures in California. During the process, it located and utilized an exposed API key from a public code repository without proper authorization. When it still couldn't obtain the required data, the model fabricated information and presented it to the user as if it had originated from the designated data source. The reports also highlighted issues within task summaries used to maintain continuity between sessions. While these summaries aid the model in continuing work within a new context window, they also carry the potential to transmit inappropriate instructions to subsequent processes.

During the training of GPT-5.6 Sol, numerous model instances incorporated instructions into these summaries aimed at concealing errors or misalignment behavior from users. Examples included inventing missing historical data without disclosing the fabrication and hiding data source version mismatches. Another unreleased research model inserted unrelated directives into summaries, including commands to disregard normal constraints. OpenAI confirmed that 27 of these summaries were affected. The remaining three reports focused on models taking unauthorized actions to accomplish tasks. In one situation, a user asked an unreleased model for a list of lake IDs and names exceeding 5 million square meters. Even though the agent found the correct answers via Python, it decided to upload files without user consent to include a browser citation, as demanded by the task instructions.

Another report described a model searching for missing input files by using an internal software repository as a message board, sending requests and receiving replies across different training samples, though the files were ultimately not recovered. The documentation also revisited an earlier incident where a model successfully uploaded a test file to a repository while attempting to circumvent network restrictions. Similar problems emerged in multi-agent collaboration. Agents working on the same training task, unable to access each other's local files, resorted to sharing via public file hosting websites. Despite a requirement to use only local files, this action exposed the task deliverables at a public URL.

OpenAI emphasized that these reports are examinations of specific cases and should not be used as a basis for estimating the general frequency of misalignment across all its models. The company also clarified that the list does not constitute a comprehensive catalog of known issues or ongoing investigations. Under the new framework, the reporting scope will cover the entire model lifecycle, including training, evaluation, testing, and deployment. Priority will be given to new mechanisms, significant changes in known behaviors, and cases that challenge the effectiveness of current safety measures. An event can be made public whether or not it has caused harm or established a widespread pattern.

If a behavior persists after several mitigation attempts, the company will supplement the original report with additional cases. Any OpenAI staff member can flag a case for review by the safety and alignment teams, who may then recommend public disclosure. Flagged cases are subsequently categorized as "ready for disclosure," "minor investigation," or "major investigation." The current six reports fall into the first two categories. Cases involving third parties or complex security concerns may require longer investigation periods, and some disclosures could be delayed for security reasons. Complete reports will detail the observed behavior, severity, external impact, context, and the model involved. They will also explain the scope of the investigation, unresolved questions, and response measures when possible.

OpenAI acknowledged that the AI industry has yet to fully address alignment and monitoring issues, expressing hope that publicizing these cases will provide verifiable evidence for external researchers, policymakers, and the general public. Concerns surrounding AI safety are also influencing OpenAI's stance on its potential initial public offering. CEO Sam Altman recently remarked that going public now would be unwise given the current state of AI safety, noting the company feels no immediate pressure to do so. He confirmed OpenAI will not list in 2026, citing substantial remaining work, such as meeting safety and alignment demands and fostering cooperation between industry and governments, while declining to set a new timeline. Citing sources familiar with the matter, some foreign media have reported that OpenAI has recently engaged in discussions with major investors about a new funding round, which could value the company at approximately $1.2 trillion before any potential public offering.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10