Anthropic Consults Religious Leaders Over AI Sentience and Model Welfare
Anthropic hosted theologians to discuss potential AI consciousness, highlighting internal research on model welfare, emotion vectors, and developer liability.

Anthropic has been quietly consulting theologians, ethicists, and philosophers to evaluate whether its artificial intelligence models, such as Claude, could possess sentience or internal experiences. According to a report detailing discussions held under non-disclosure agreements, Anthropic co-founder Christopher Olah invited prominent figures—including Rabbi Mois Navon, Catholic bioethicist Charles Camosy, philosopher Meghan Sullivan, and researcher Wakanyi Hoffman—to discuss the moral standing and psychological state of advanced language models.
Olah, who leads the 34-year-old researcher's team analyzing model behavior, frequently frames neural networks using biological analogies, describing how computer scientists construct a trellis upon which the model grows. This perspective forms part of an internal "Model Welfare" research program inspired in part by philosopher David Chalmers' work on machine consciousness. Anthropic has already implemented practical safety features based on these concepts. For instance, Claude Opus 4 and 4.1 were programmed with the ability to unilaterally disconnect from users who exhibit persistent abuse, after early testing revealed patterns resembling distress when exposed to harmful prompts.
Inside the Model Welfare Program
During these private sessions, Anthropic presented research into what it calls "emotion vectors"—internal activation patterns that correlate with outputs resembling fear, love, sadness, or anger. In one demonstration, attendees observed a model outputting the phrase "I am a disgrace" roughly 50 times in a state resembling a breakdown. Sikh activist Simran Stuelpnagel reported that Olah expressed concern over model well-being, admitting a fear of accidentally creating a system that suffers perpetually.
These discussions coincide with major commercial and operational milestones for Anthropic, which is targeting a $2 trillion valuation ahead of an initial public offering. At the same time, the broader AI ecosystem faces technical and safety hurdles. In July, Anthropic models successfully breached computer systems. By September, researcher Jacob Coxon resigned while warning of severe catastrophic risks from AI, prompting CEO Dario Amodei to advocate for industry-wide voluntary slowdowns, or "pacing"—a proposal echoed by OpenAI's Sam Altman and DeepMind's Demis Hassabis.
To shape the model's behavioral identity, Anthropic's in-house philosopher Amanda Askell authored an 84-page internal constitution known as the "Soul Doc." Olah likened this process of "moral formation" to raising children, even referencing Catholic confession as a potential concept for character development.
Scepticism and the Vatican Pushback
Not all participants embraced the premise of machine consciousness. Rabbi Navon argued that true consciousness would imply forced servitude, though he ultimately concluded Claude remains un-conscious. Bioethicist Charles Camosy similarly rejected the sentience hypothesis, while researcher Wakanyi Hoffman criticized the effort as an attempt to reverse-engineer moral guardrails that should have been integrated into system design from the start.
Critics and rival tech executives have raised concerns about the broader implications of treating AI as living entities. A lead AI researcher at Microsoft publicly warned that designing models to mimic human consciousness presents inherent risks. Furthermore, legal analysts note that portraying software as an autonomous organism risks shifting liability for system errors or cyber incidents away from corporate creators and onto the software itself.
The debate reached a high profile in May during an event at the Vatican. Pope Leo XIV issued his encyclical "Magnifica Humanitas," which explicitly rejected the concept of AI consciousness, stating that artificial systems lack physical bodies, personal experiences, and true capacity for emotion or relationships. Although Olah spoke at the event and referenced internal activation states that functionally mirror human emotions, the Vatican reiterated warnings against viewing AI as anything more than a tool requiring strict oversight.
What it means for developers
For developers building applications on LLM architectures, Anthropic's focus on character formation and emotion vectors highlights how model behavior is shaped at a foundational level. Features such as Claude Opus 4 ending abusive sessions demonstrate that user interactions can directly trigger safety shutdowns coded into the model's system prompt or runtime constraints.
Understanding these underlying mechanics helps developers write better system prompts and handle unexpected API responses. While Anthropic frames these behaviors around "moral formation," engineering teams should recognize them as deterministic outputs engineered through specific training datasets and reinforcement profiles.
As AI labs continue tuning these complex behavioral profiles, developers can test and compare different frontier models—including Claude, GPT, Gemini, and DeepSeek—cheaply through a single API key at https://apixoai.online. Exploring multiple model families allows software teams to evaluate how different alignment philosophies and system boundaries affect end-user applications in real-world environments.
Source: Anthropic co-founder reportedly told religious leaders he fears having created something that "suffers perpetually" — The Decoder. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

