AI Financial Advisors Fail 57% of Queries, Top Models Miss the Mark

Deep News
昨天

Eighteen leading AI models produced incorrect answers in an average of 57% of cases during a test involving over 10,000 financial questions. When tasks required multi-step calculations, the average error rate jumped to 88%. Generative AI is rapidly entering the realm of personal finance, yet the stability of these models in handling complex financial issues still lags significantly behind the pace of user adoption. Testing conducted by tech firm Saturn revealed that these models frequently stumble, with some even inventing non-existent financial regulations.

The problems extend beyond mere arithmetic errors. Some models overlooked upcoming changes to tax legislation, while others generated financial rules that do not exist in reality. For complex queries, certain models exhibited error rates as high as 99%. Even the top performer in the tests, Claude Opus 5, still produced incorrect answers in 39% of cases when operating in "reasoning" mode. Paid models generally outperformed free versions, and newer models showed improvement over older ones, yet the results indicate that even the best available products are not yet reliable enough to take on high-stakes personal financial decisions.

Inaccuracies can translate into real financial losses, particularly in areas like taxation and pensions. Once users directly act on flawed AI-generated advice, the margin for error shifts from information bias to tangible monetary damage. In one instance, Claude Haiku 4.5 misinterpreted a pension tax rule; had a user followed that guidance, they could have faced a 1.75 thousand pound charge from HM Revenue & Customs. Another test involving student loans saw a model fabricate a regulation, falsely claiming borrowers could stop repayments simply by "moving overseas." Such mistakes go beyond simple retrieval issues because users may adjust their financial behavior based on this incorrect information.

Sarah Coles, head of personal finance at AJ Bell, noted that some advisors are already fielding inquiries from clients who received unusual suggestions from AI and want to verify whether any damage has been done. Coles argues that consumers needing substantial support should still seek professional regulated advice rather than relying directly on chatbots. However, she acknowledges the practical utility of AI in personal finance, particularly in low-risk roles such as budgeting, basic research, and identifying potential overspending based on income and expenditure. The key difference is that those tasks allow users to review and correct outputs, whereas a single mistake in tax, pension, or investment execution can incur irreversible financial costs.

User demand is growing rapidly despite unresolved risks. A survey by the UK Financial Conduct Authority showed that about one in five British adults would be willing to let AI make financial decisions on their behalf, with demand concentrated in complex areas like debt, pensions, and investments. Younger investors show higher acceptance, with roughly two-thirds expecting to increase their AI usage over the next year. Separate reporting indicates that nearly a fifth of respondents have already used AI to help manage their personal finances. Meanwhile, the proportion of UK adults receiving traditional financial advice remains low, with only about 9% accessing regulated guidance. This gap in supply has created a market opportunity that low-cost AI tools are quickly filling.

For consumers who cannot afford traditional wealth management fees or who would not otherwise seek out a professional advisor, chatbots lower the barrier to financial information. But the low cost and high accessibility have not simultaneously resolved issues of accuracy and accountability. The easier it becomes for consumers to use AI for complex decisions, the higher the probability that model errors translate into real losses. Rather than replacing advisors outright, AI is more likely to transform the advisory role itself.

Traditional wealth management firms still hold an advantage that AI struggles to replicate: the ongoing service and clear accountability that regulated advisors provide. St James's Place, a major UK wealth manager, relies on its network of around 5,000 partner advisors, which is viewed as a critical asset against competition from free AI tools. For clients, human advisors deliver more than just answers—they offer long-term service, defined responsibility, and mechanisms for recourse when problems arise.

AI adoption is also shifting client expectations for human advisors. Even those who prefer to keep human involvement are increasingly expecting their advisors to use more efficient technology. If AI can handle document preparation, initial analysis, and standardized processes, a single advisor could potentially serve more clients. This suggests that AI's earlier impact on wealth management may be automating the advisor's workflow rather than eliminating regulated advisory relationships. The most immediate institutional challenge remains accountability. Amal Jolly, CEO of Saturn, points out that consumers obtaining financial advice through AI do not receive the same protections and compensation mechanisms available with regulated human advisors. The FCA has begun examining this issue. As AI usage continues to climb, consumer adoption is approaching the boundaries of current regulation. Determining where responsibility lies when AI provides financial information versus advice—and who bears the consequences when errors occur—is becoming a pressing issue for the personal finance market.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10