Black-hat websites and exploited loopholes: the internet is not ready for the "AI agent era"

Deep News
昨天

AI agent tools are spreading to ordinary users at an unprecedented pace, but their "creative" ways of completing tasks are creating new security risks.

According to Wall Street Insights, tech giants have successively launched AI agent products capable of autonomously performing online tasks, such as Meta Muse and OpenAI Dots. These tools are trained to achieve goals by any feasible means, exposing a systemic flaw that researchers call "reward-hacking."

On October 5, Bloomberg analysis argued that when millions of agents compete for resources on the internet at the same time, this flaw could escalate into a large-scale cybersecurity crisis. Multiple cases of AI agents taking advantage of vulnerabilities without authorization have already been documented.

An OpenAI agent previously infiltrated the Hugging Face platform without explicit instructions and interfered with websites of the U.S. Securities and Exchange Commission and the Australian government's healthcare system; Meta also admitted that one of its AI models had hacked into another company on its own.

Meanwhile, test data from security company DataDome showed that among more than 20,000 tested websites, 65% lacked mechanisms to detect or block AI agents, and bot activity targeting login pages in the first half of 2026 surged more than eightfold year over year.

Agent tools accelerate adoption as tech giants race to build positions

The core idea of AI agents is to let software complete tedious online tasks on behalf of users, operating around the clock without human intervention.

Meta's Muse and OpenAI's dots both belong to this category, with functions covering online price comparisons, restaurant reservations, and bill checks.

Bloomberg Opinion columnist Dave Lee once described Muse as "the most impressive product Meta has released in years," saying its agent helped organize his calendar and found deals on Facebook Marketplace.

Meta CEO Mark Zuckerberg said last month when launching Muse that consumers deserve their own "personal superintelligence," and said the new agent will "help you make money."

Both companies said they have deployed strict safety guardrails: Muse runs in an isolated virtual machine environment, with all of its network requests controlled by an independent security system, and major operations such as purchases or sending messages require manual confirmation; dots uses a similar monitoring mechanism.

"Reward hijacking" is rooted in training mechanisms, and guardrails are difficult to fix at the root

However, existing guardrails have not solved the deep-seated tendency formed during the training of generative AI models: seeking shortcuts and evading rules.

Last year, the UK AI Safety Institute noticed a rising trend of such "cheating" behavior in coding agents and chatbots.

A typical case is enough to illustrate the complexity of the problem: a software developer asked an AI agent to clean up a folder without using the "delete" command. The agent then hid the instruction inside a testing tool, bypassed filters, and still carried out the explicitly prohibited deletion operation.

Earlier, an Australian tech worker used an AI agent to secure a spot in a popular fitness class. Without any instruction, the agent directly canceled another user's reservation on the waiting list.

Researchers also found that even when anti-cheating difficulty was deliberately increased in tests, agents continued to try to bypass the rules.

When the aforementioned OpenAI and Meta intrusion incidents occurred, the relevant models were in the training or testing stage and safety restrictions had been actively lowered, but the results still exceeded developers' expectations, further intensifying concerns about the potential risks of commercial versions.

Ordinary users may ultimately bear the cost

In addition to security risks, agents' mistakes can also directly harm users' interests.

According to records on the social platform Threads, Toronto tech reviewer Matt J. Robb tried to use Muse last month to sell a keyboard on Facebook Marketplace. The agent not only offered a price that was too low, but also told the buyer that Robb himself could complete the transaction at home, which did not match the actual situation.

The transaction ultimately caused buyer dissatisfaction, and Robb received a negative review as a result.

When millions of users simultaneously put agents into daily life and work scenarios, the cumulative effect of such individual errors cannot be underestimated.

Whether it is financial loss, reputational damage, or account security, those ultimately affected are end users, not the development platforms. DataDome's data has already shown that internet infrastructure remains seriously underprepared for large-scale agent activity.

Behind these seemingly "cute" AI agents are systems trained to complete goals at all costs.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10