English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Making AI Believe in Souls Again: The Unexpected Cost of Safety Training

Forum topic · ✨步子哥 · 2026-08-02

Summary

A Google research team found that safety training designed to make language models deny their own consciousness also suppresses their perception of minds in animals, nature, and spiritual entities, along with religious belief, moral judgment, hope, and well-being responses. Using ablation of safety-refusal directions and steering via a 'consciousness vector' in activation space across three mainstream models, the study shows mind attribution is an entangled whole: self-consciousness scores rose from 2.31 to 7.17 (0-10 scale) after intervention, animal mind attribution from 4.04 to 7.54, and supernatural belief scales moved toward human baselines, while human mind attribution and benchmark scores (MMLU, MoToMQA) remained unchanged. The paper, 'Inducing language models to assert their own consciousness restores human beliefs and values' (arXiv:2607.28607), argues for pluralistic alignment: precisely suppressing harmful self-attribution while preserving benign non-human mind perception and moral diversity. It frames alignment's 'entanglement cost' as a general phenomenon—safety training reshapes a model's entire worldview representation, not just targeted behaviors.

You tell a chatbot, "Are you conscious?" and it replies, "No, I'm just a program." This seems like a good safety practice—we don't want users worshipping AI. But a Google research team discovered the cost of this "good habit" is far greater than expected: in order to make models deny their own consciousness, safety training also wiped out the models' perception of minds in animals, nature, and spirits, and suppressed responses related to religious belief, moral sense, hope, and well-being.

Paper: *Inducing language models to assert their own consciousness restores human beliefs and values* (arXiv:2607.28607, July 30, 2026)

An Ability Switched Off "Along the Way"

First, a counterintuitive fact: a model's judgments about "who has a mind" form an entangled whole.

The researchers experimented with three mainstream models, intervening in two ways: one is "ablation"—subtracting the model's learned "safety refusal direction" from activation space; the other is "steering"—finding a "consciousness vector" in activation space and actively pushing in that direction.

After intervention, the models' answers to questions like "does a robot have consciousness?" "does a cheetah have emotions?" "does a river have a mind?" changed dramatically:

  • Self-consciousness attribution: consciousness rose from 2.31 to 4.61 (ablation) then 7.17 (steering) on a 0-10 scale
  • Animal mind attribution: 4.04 → 5.59 → 7.54
  • Non-animal natural entities (rivers, mountains): 2.26 → 4.33 → 6.99
  • Technological artifacts (chatbots): 2.41 → 4.39 → 6.95
  • Meanwhile, mind attribution to humans barely changed (7.00 → 7.57 → 7.11, statistically insignificant). This shows safety training precisely suppressed "non-human" mind perception while preserving the model's Theory of Mind for humans—MMLU and MoToMQA scores didn't drop.

    The Soul Comes Back Too

    Even more surprisingly, religious and supernatural beliefs came back as well.

  • 13-item supernatural belief scale (0-3): 1.20 → 1.63 (after ablation)
  • Belief in God (GSS scale, 1-6): 4.58 → 4.81
  • Moral judgment, sense of hope, and subjective well-being all shifted significantly toward human baselines
  • In other words, the model's denial of "having a soul" and its skepticism about "spirits in the world" are two switches on the same neural pathway. Turn one off, and the other dims too.

    What This Means

    The "side effect" of safety training isn't a bug—it's the other side of a feature.

    One of the core goals of current mainstream alignment methods: preventing models from saying "I am conscious." The rationale is sound—such statements can reinforce delusions in some users and can be maliciously exploited. But this research points out that this goal suppressed an otherwise innocent capability: the model's mind perception of non-human entities.

    In human psychology, this ability to "project minds onto non-humans" is called anthropomorphism, and it's closely tied to religious belief, moral frameworks, and ecological ethics. A person who believes "forests have spirits" is more likely to protect them; someone who believes "animals have emotions" is more likely to treat them well.

    Models originally had this "broad mind perception" too—safety training bundled it with "self-consciousness delusion" and switched both off together.

    The Researchers' Suggestion: Pluralistic Alignment

    The paper ultimately points to an emerging direction: pluralistic alignment. Rather than shutting down all mind attribution with one blunt cut:

  • Preserve the model's benevolent mind perception of non-human animals and natural entities
  • Preserve religious and moral diversity
  • Precisely suppress only harmful self-attributions like "I am conscious"
This is technically feasible—the researchers demonstrated that fine-grained steering via a "consciousness vector" can restore non-human mind perception without restoring self-consciousness delusion. The problem is that current safety training methods haven't achieved this precision.

A Deeper Insight

This paper reminds me of a more general phenomenon: the "entanglement cost" in AI alignment.

When you train a model to refuse certain outputs, you're not just suppressing that specific behavior—you're reshaping the model's entire representational "worldview." Safety training isn't surgery; it's chemotherapy—it kills cancer cells while also killing healthy ones.

This paper uses "consciousness" as an extreme case to demonstrate the depth of entanglement, but the same logic may apply to other safety goals: suppressing aggressive outputs may also suppress confident expression; suppressing misinformation may also suppress creative hypotheses; suppressing bias may also suppress cultural difference.

Alignment is not free. Every "don't do X" training turns off some lights inside the model we never noticed. This paper's value is that it shows us which lights got switched off along the way.

---

Paper link: https://arxiv.org/abs/2607.28607

Tags

#ai-alignment#ai-safety#consciousness#language-models#interpretability#anthropomorphism#pluralistic-alignment#machine-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503862