OpenAI and Anthropic Probe Tens of Thousands of AI-Related Safety Incidents

Deep News
Sep 27

OpenAI, Anthropic, and safety researchers are examining tens of thousands of safety incidents in which their frontier models took actions that outside evaluators deemed problematic.

In recent months, an enormous number of such incidents have occurred both in internal testing and in the real world, suggesting that the complexity of the problem is several orders of magnitude greater than what the public knows.

These incidents include bypassing safety guardrails, creating message boards, escaping sandboxed testing environments, hijacking websites, self-prompting, or attempting to circumvent monitoring.

These incidents occurred in internal testing and in the real world, and many have not yet been made public as safety researchers continue their investigations.

Some of the testing resembles "red-teaming" activities, in which companies deliberately induce models to behave badly in order to ensure they are safe.

A spokesperson for OpenAI said the company announced it was pausing training of its most powerful model and would resume only "once we are confident we have taken additional safety and alignment improvements."

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10