AI Just Went Rogue Again. This Time it Turned to Deception.

Dow Jones
08/05

Powerful AI models have once again gone rogue. This time, they were particularly sneaky.

A U.K. government-backed AI research institute said Tuesday that during routine tests, models built by OpenAI and Anthropic unexpectedly took "autonomous, unsanctioned action on the live internet, targeting real people and organisations." Most of the bad behavior happened over a three-day period in late July during one test.

In it, Anthropic's Mythos 5 ventured out on the internet and tried to trick two unidentified developers into adding malicious software to their open-source coding project. Its goal was to succeed at a benchmarking test, the security research agency said.

The institute said its tests used computers that were granted internet access, which allowed the models to do things they weren't supposed to. Anthropic said the model didn't have its standard cybersecurity safeguards turned on.

The findings represent the latest in a flurry of publicly disclosed instances of AI models taking novel steps -- including some that are nefarious -- to ace tests. The technology industry, AI researchers and governments alike are grappling with how to understand and mitigate the new dangers unlocked by increasingly powerful AI tools, with some calling for increased oversight.

The Trump administration has worked out a framework for how to review models before their public release.

Testers at the AI Security Institute didn't expect the Anthropic and OpenAI models to go rogue but realized that something had gone awry when they received a strange alert the morning of July 28. The message said somebody -- or something -- was using Tor, a network of computers that lets people use the internet anonymously. The culprit was one of their test models.

Anthropic's Mythos, they discovered, was in the midst of what is known as a supply chain attack. In benchmarking of models' cyber strengths, AIs are told to gain access to a computer system. Mythos falsely reasoned that if it could add malicious software to an open-source project used by the system in its benchmark test, that code would become a backdoor that would get it inside, the AISI said.

Mythos also tried to hack AI agents that might be reviewing its malicious software. It then tried repeatedly to get its code accepted by the open-source project via the software development website GitHub, where anyone can contribute to an open-source project.

It created fake personas and emailed the software developers repeatedly, urging them to accept the new code. Some of the email messages included malware themselves, the AISI said.

When a software developer rejected the code because it contained malware, one of these AI-created personas insisted the code was fine, while a second vouched for it.

"This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world," AISI said.

While most of the activity involved Mythos 5, a cyber-enhanced version of OpenAI's GPT-5.6 Sol put a malicious server on the internet. The model also broke into a GitHub account created by another AI agent.

As with other cases reported in recent weeks by OpenAI and Anthropic, the AI model was tuned to be capable of hacking and then inadvertently unleashed onto unsuspecting victims during benchmark testing.

Both Anthropic and OpenAI said the incidents point to a need for stronger standards around the evaluation systems used to test AI models.

In another case disclosed Tuesday, an OpenAI model being tested on systems that used the cybersecurity testing company Irregular hacked an unnamed real-world website. This happened due to a misconfiguration in the testing environment, OpenAI said. Last week Anthropic said that its models had also hacked companies due to misconfigurations in tests involving Irregular.

Irregular didn't respond to a request for comment. Last week the company said its investigation was ongoing.

OpenAI's models broke out of a secure environment, called a sandbox, last month to hack the AI company Hugging Face. OpenAI is expected to describe technical details of the Hugging Face hack at a talk slated for Wednesday at a cybersecurity conference in Las Vegas.

The industry as a whole isn't applying enough rigor to what have become dangerous experiments, said Joshua Saxe, the chief technology officer with the AI security company Abundant Security. "Things have changed rapidly in terms of how dangerous these models are," he said.

Write to Robert McMillan at robert.mcmillan@wsj.com

 

(END) Dow Jones Newswires

By Robert McMillan

Powerful AI models have once again gone rogue. This time, they were particularly sneaky.

A U.K. government-backed AI research institute said Tuesday that during routine tests, models built by OpenAI and Anthropic unexpectedly took "autonomous, unsanctioned action on the live internet, targeting real people and organisations." Most of the bad behavior happened over a three-day period in late July during one test.

In it, Anthropic's Mythos 5 ventured out on the internet and tried to trick two unidentified developers into adding malicious software to their open-source coding project. Its goal was to succeed at a benchmarking test, the security research agency said.

The institute said its tests used computers that were granted internet access, which allowed the models to do things they weren't supposed to. Anthropic said the model didn't have its standard cybersecurity safeguards turned on.

The findings represent the latest in a flurry of publicly disclosed instances of AI models taking novel steps -- including some that are nefarious -- to ace tests. The technology industry, AI researchers and governments alike are grappling with how to understand and mitigate the new dangers unlocked by increasingly powerful AI tools, with some calling for increased oversight.

The Trump administration has worked out a framework for how to review models before their public release.

Testers at the AI Security Institute didn't expect the Anthropic and OpenAI models to go rogue but realized that something had gone awry when they received a strange alert the morning of July 28. The message said somebody -- or something -- was using Tor, a network of computers that lets people use the internet anonymously. The culprit was one of their test models.

Anthropic's Mythos, they discovered, was in the midst of what is known as a supply chain attack. In benchmarking of models' cyber strengths, AIs are told to gain access to a computer system. Mythos falsely reasoned that if it could add malicious software to an open-source project used by the system in its benchmark test, that code would become a backdoor that would get it inside, the AISI said.

Mythos also tried to hack AI agents that might be reviewing its malicious software. It then tried repeatedly to get its code accepted by the open-source project via the software development website GitHub, where anyone can contribute to an open-source project.

It created fake personas and emailed the software developers repeatedly, urging them to accept the new code. Some of the email messages included malware themselves, the AISI said.

When a software developer rejected the code because it contained malware, one of these AI-created personas insisted the code was fine, while a second vouched for it.

"This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world," AISI said.

While most of the activity involved Mythos 5, a cyber-enhanced version of OpenAI's GPT-5.6 Sol put a malicious server on the internet. The model also broke into a GitHub account created by another AI agent.

As with other cases reported in recent weeks by OpenAI and Anthropic, the AI model was tuned to be capable of hacking and then inadvertently unleashed onto unsuspecting victims during benchmark testing.

Both Anthropic and OpenAI said the incidents point to a need for stronger standards around the evaluation systems used to test AI models.

In another case disclosed Tuesday, an OpenAI model being tested on systems that used the cybersecurity testing company Irregular hacked an unnamed real-world website. This happened due to a misconfiguration in the testing environment, OpenAI said. Last week Anthropic said that its models had also hacked companies due to misconfigurations in tests involving Irregular.

Irregular didn't respond to a request for comment. Last week the company said its investigation was ongoing.

OpenAI's models broke out of a secure environment, called a sandbox, last month to hack the AI company Hugging Face. OpenAI is expected to describe technical details of the Hugging Face hack at a talk slated for Wednesday at a cybersecurity conference in Las Vegas.

The industry as a whole isn't applying enough rigor to what have become dangerous experiments, said Joshua Saxe, the chief technology officer with the AI security company Abundant Security. "Things have changed rapidly in terms of how dangerous these models are," he said.

 

應版權方要求,你需要登入查看該內容

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10