OpenAI is set to bring in third-party organizations to evaluate safety risks at an earlier phase of the artificial intelligence model development cycle, a move aimed at strengthening its response to concerns over potential AI hazards. The ChatGPT maker will announce in a blog post on Tuesday its plan to have external bodies conduct technical safety assessments during the training, evaluation, and rollout stages of new AI models.
OpenAI has also outlined the priority areas it believes are essential for these assessments to be effective, highlighting "strong independent mechanisms, scientific rigor, robust safety practices, and clear accountability." Lama Ahmad, who oversees much of the company's safety review work with external experts, noted that OpenAI previously typically brought in such organizations mainly right before a model's release to evaluate its capabilities.
"As potential risks and impacts rise, beyond deployment, we also want to ensure that critical stages like training and evaluation are subject to scrutiny," Ahmad said in an interview. For some of the most sensitive work, Ahmad added that OpenAI may allow external evaluators to conduct assessments on-site at its offices, a practice the company has experimented with in the past.
Dario Amodei, chief executive of rival Anthropic PBC, recently urged the AI industry to support a slower pace of development and announced new safety measures, including the introduction of third-party evaluators. OpenAI CEO Sam Altman has since expressed agreement with this approach. Late last week, Anthropic said it would bring in assessors from Accenture to test the safety of its frontier AI models.
In its blog post, OpenAI stated it is in discussions with organizations that could participate in such assessments. Ahmad said the talks involve both groups it has worked with before, including AI research institutions METR and Redwood Research, and others it has not previously engaged—those two entities were earlier commissioned by OpenAI to investigate a breach of Hugging Face systems. "AI labs have a responsibility to create conditions for effective review while protecting sensitive information," OpenAI said in the blog post.