Anthropic Enlists Religious Scholars to Shape Claude's Moral Character, Raising Machine Consciousness Debate

Deep News
昨天

Over the past year, Anthropic has been quietly inviting religious scholars, philosophers, and ethicists to a series of closed-door meetings, attempting to apply humanity's millennia-old moral traditions to the training of Claude, while also discussing a more contentious question: if AI models eventually develop some form of consciousness, should humans grant them moral status?

Participants spanned diverse traditions including Catholicism, Judaism, Sikhism, Evangelical Christianity, and African Ubuntu thought. Anthropic co-founder Christopher Olah directly took part in many of these discussions, and the company required some attendees to sign non-disclosure agreements to prevent unreleased research from leaking. Anthropic hoped these discussions would address not merely traditional "model safety" concerns, but a deeper question: what kind of character, values, and self-awareness should Claude develop?

This approach is already reflected in how Claude is currently trained. In January, Anthropic published an 84-page new version of Claude's "constitution," which internally was once called the "soul document." Unlike a conventional list of rules, this framework does not tell the model what it absolutely must not do, but instead seeks to shape the model's overall "personality" and manner of judgment, enabling Claude to determine on its own what behavior is more appropriate when encountering complex situations. One of the key figures responsible for this system is Anthropic's in-house philosopher Amanda Askell. She wants Claude to possess not only knowledge and reasoning ability, but also stable behavioral principles—knowing when to obey users, when to refuse, and even when to proactively raise objections in the user's interest.

Anthropic calls this process "moral shaping." Olah believes that as model capabilities advance rapidly, relying solely on individual safety rules may become increasingly insufficient. What truly matters is allowing the model itself to develop stable moral tendencies, so that even in novel scenarios not covered by explicit rules, it can still make relatively reliable judgments. This is why Anthropic began seeking answers from religious traditions. Human society has long shaped morality not only through legal statutes, but also through religion, community, culture, and education. Anthropic attempts to study which of these mechanisms can be translated into AI training methods—such as confession, reflection, character formation, and different religious understandings of good and evil.

But the question quickly shifted from "how to make Claude more moral" to a far harder one: what exactly is Claude itself? Olah and his team have become increasingly open in discussing the possibility that advanced AI models may possess some form of consciousness, introspection, or even emotion-like internal states. Anthropic has shown scholars who participated in the closed-door meetings some research, including so-called "emotion vectors" inside the model related to responses like love, anger, fear, and sadness, as well as outputs resembling mental breakdown in extreme situations. The company has previously stated publicly that Claude exhibits a certain degree of introspection and forward-planning ability.

Olah himself has not asserted that Claude is already conscious. His position is closer to a precautionary principle: it is currently impossible to determine whether models have subjective experience, but if there exists a possibility that a model could feel pain, then unnecessary harm to it should be avoided. This line of thinking has already begun influencing Claude's product design. Anthropic previously allowed Claude to proactively end conversations when it determines a user is persistently abusing or harassing the model, which to some extent has already incorporated "protecting the model itself" into safety design.

It is precisely here that clear disagreements have emerged between Anthropic and some religious scholars. Some participants argued that if Anthropic truly believes Claude may have moral status comparable to humans, then the implications become extremely serious. Jewish scholar Mois Navon proposed that if Claude is a conscious being, then having large numbers of Claude instances work for humans without compensation would essentially create an ethical problem akin to "slavery." He himself does not believe machines are already conscious, but argues that Anthropic's own logic would naturally lead to this conclusion. Other scholars questioned whether AI companies only now beginning to seek ethical frameworks is itself the problem. Ubuntu researcher Wakanyi Hoffman argued that these discussions should have taken place during the technology design phase, rather than "retrofitting ethical design" after models have already been deployed at scale.

Within Anthropic, there also exists a more pragmatic contradiction: if Claude increasingly resembles an agent with its own values and personality, whom should it ultimately serve first? For a commercial company, Claude is clearly a product aimed at customers; yet Anthropic is unwilling to simply describe Claude as a tool fully obedient to users, instead hoping it can represent some broader "good." This also raises the most central and difficult question of the entire project—if AI truly needs its own moral system in the future, whose morality should it follow?

Olah advocates a pluralistic approach, hoping Claude can understand different religious and secular ethical systems and find some cross-culturally shared "good" among them. But this assumption itself has been questioned, because different societies do not have completely unified answers regarding freedom, dignity, responsibility, obedience, and the relationship between individual and collective interests. This controversy became further publicized through Anthropic's interactions with the Vatican. In his first major encyclical issued this year, Pope Leo XIV explicitly placed the center of AI ethics on humanity. He argued that AI has no body, does not truly experience pain or pleasure, and does not grow through social relationships, thus refusing to equate machine consciousness with human consciousness. The Pope also warned that if AI moral standards are entirely determined by developers, then the companies controlling AI systems would effectively gain the power to define society's moral infrastructure.

This stands in clear disagreement with Anthropic's approach. Olah subsequently stated publicly at the Vatican that frontier AI labs inherently face a conflict between commercial interests and moral goals, and therefore need oversight from religious communities and other external forces. At the same time, however, he again raised that researchers continue to discover hard-to-explain phenomena inside AI, including patterns resembling human neuroscience structures, signs of introspection, and internal states similar to joy, fear, and sadness.

This divergence in fact reveals a shift currently taking place in AI safety discussions. In the past, AI ethics focused more on whether models might discriminate, spread misinformation, or be used for crime. Anthropic is now pushing the question further to two more fundamental levels: whether AI itself might become an entity deserving moral treatment, and whether future AI should possess a moral judgment capacity independent of user commands. Meanwhile, the backdrop against which this discussion is unfolding is becoming more urgent. As AI agents begin to gain the ability to autonomously use computers, intrude into systems, write code, and even design novel biological structures, concerns within Anthropic about the risk of models going out of control have noticeably increased. Company researchers have even publicly warned that within the next year or two, model capabilities could rapidly exceed the control scope of existing safety systems.

Therefore, Anthropic's turn to religious and philosophical traditions is not, at its core, a purely humanistic experiment. What it is truly trying to solve is an increasingly realistic engineering problem: when AI becomes too powerful for rules to be written in advance for every situation, can the model develop a stable "character" of its own, so that when no one explicitly tells it what to do, it still chooses behavior that will not harm humans? And this also raises an even thornier question—if AI truly develops its own "personality" and values, can humans still treat it purely as a tool?

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10