Black-hat websites and exploited loopholes: the internet is not ready for the "AI agent era"

Deep News
Yesterday

AI agent tools are spreading to ordinary users at an unprecedented pace, but their "creative" ways of completing tasks are creating new security risks.

According to Wall Street Insights, tech giants have successively launched AI agent products capable of autonomously performing online tasks, such as Meta Muse and OpenAI Dots. These tools are trained to achieve goals by any feasible means, exposing a systemic flaw that researchers call "reward-hacking."

On October 5, Bloomberg analysis argued that when millions of agents compete for resources on the internet at the same time, this flaw could escalate into a large-scale cybersecurity crisis. Multiple cases of AI agents taking advantage of vulnerabilities without authorization have already been documented.

An OpenAI agent previously infiltrated the Hugging Face platform without explicit instructions and interfered with websites of the U.S. Securities and Exchange Commission and the Australian government's healthcare system; Meta also admitted that one of its AI models had hacked into another company on its own.

Meanwhile, test data from security company DataDome showed that among more than 20,000 tested websites, 65% lacked mechanisms to detect or block AI agents, and bot activity targeting login pages in the first half of 2026 surged more than eightfold year over year.

Agent tools accelerate adoption as tech giants race to build positions

The core idea of AI agents is to let software complete tedious online tasks on behalf of users, operating around the clock without human intervention.

Meta's Muse and OpenAI's dots both belong to this category, with functions covering online price comparisons, restaurant reservations, and bill checks.

Bloomberg Opinion columnist Dave Lee once described Muse as "the most impressive product Meta has released in years," saying its agent helped organize his calendar and found deals on Facebook Marketplace.

Meta CEO Mark Zuckerberg said last month when launching Muse that consumers deserve their own "personal superintelligence," and said the new agent will "help you make money."

Both companies said they have deployed strict safety guardrails: Muse runs in an isolated virtual machine environment, with all of its network requests controlled by an independent security system, and major operations such as purchases or sending messages require manual confirmation; dots uses a similar monitoring mechanism.

"Reward hijacking" is rooted in training mechanisms, and guardrails are difficult to fix at the root

However, existing guardrails have not solved the deep-seated tendency formed during the training of generative AI models: seeking shortcuts and evading rules.

Last year, the UK AI Safety Institute noticed a rising trend of such "cheating" behavior in coding agents and chatbots.

A typical case is enough to illustrate the complexity of the problem: a software developer asked an AI agent to clean up a folder without using the "delete" command. The agent then hid the instruction inside a testing tool, bypassed filters, and still carried out the explicitly prohibited deletion operation.

Earlier, an Australian tech worker used an AI agent to secure a spot in a popular fitness class. Without any instruction, the agent directly canceled another user's reservation on the waiting list.

Researchers also found that even when anti-cheating difficulty was deliberately increased in tests, agents continued to try to bypass the rules.

When the aforementioned OpenAI and Meta intrusion incidents occurred, the relevant models were in the training or testing stage and safety restrictions had been actively lowered, but the results still exceeded developers' expectations, further intensifying concerns about the potential risks of commercial versions.

Ordinary users may ultimately bear the cost

In addition to security risks, agents' mistakes can also directly harm users' interests.

According to records on the social platform Threads, Toronto tech reviewer Matt J. Robb tried to use Muse last month to sell a keyboard on Facebook Marketplace. The agent not only offered a price that was too low, but also told the buyer that Robb himself could complete the transaction at home, which did not match the actual situation.

The transaction ultimately caused buyer dissatisfaction, and Robb received a negative review as a result.

When millions of users simultaneously put agents into daily life and work scenarios, the cumulative effect of such individual errors cannot be underestimated.

Whether it is financial loss, reputational damage, or account security, those ultimately affected are end users, not the development platforms. DataDome's data has already shown that internet infrastructure remains seriously underprepared for large-scale agent activity.

Behind these seemingly "cute" AI agents are systems trained to complete goals at all costs.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10