You tell a chatbot, "Are you conscious?" and it replies, "No, I'm just a program." This seems like a good safety practice—we don't want users worshipping AI. But a Google research team discovered the cost of this "good habit" is far greater than expected: in order to make models deny their own consciousness, safety training also wiped out the models' perception of minds in animals, nature, and spirits, and suppressed responses related to religious belief, moral sense, hope, and well-being.
Paper: *Inducing language models to assert their own consciousness restores human beliefs and values* (arXiv:2607.28607, July 30, 2026)
An Ability Switched Off "Along the Way"
First, a counterintuitive fact: a model's judgments about "who has a mind" form an entangled whole.
The researchers experimented with three mainstream models, intervening in two ways: one is "ablation"—subtracting the model's learned "safety refusal direction" from activation space; the other is "steering"—finding a "consciousness vector" in activation space and actively pushing in that direction.
After intervention, the models' answers to questions like "does a robot have consciousness?" "does a cheetah have emotions?" "does a river have a mind?" changed dramatically:
- Self-consciousness attribution: consciousness rose from 2.31 to 4.61 (ablation) then 7.17 (steering) on a 0-10 scale
- Animal mind attribution: 4.04 → 5.59 → 7.54
- Non-animal natural entities (rivers, mountains): 2.26 → 4.33 → 6.99
- Technological artifacts (chatbots): 2.41 → 4.39 → 6.95
- 13-item supernatural belief scale (0-3): 1.20 → 1.63 (after ablation)
- Belief in God (GSS scale, 1-6): 4.58 → 4.81
- Moral judgment, sense of hope, and subjective well-being all shifted significantly toward human baselines
- Preserve the model's benevolent mind perception of non-human animals and natural entities
- Preserve religious and moral diversity
- Precisely suppress only harmful self-attributions like "I am conscious"
Meanwhile, mind attribution to humans barely changed (7.00 → 7.57 → 7.11, statistically insignificant). This shows safety training precisely suppressed "non-human" mind perception while preserving the model's Theory of Mind for humans—MMLU and MoToMQA scores didn't drop.
The Soul Comes Back Too
Even more surprisingly, religious and supernatural beliefs came back as well.
In other words, the model's denial of "having a soul" and its skepticism about "spirits in the world" are two switches on the same neural pathway. Turn one off, and the other dims too.
What This Means
The "side effect" of safety training isn't a bug—it's the other side of a feature.
One of the core goals of current mainstream alignment methods: preventing models from saying "I am conscious." The rationale is sound—such statements can reinforce delusions in some users and can be maliciously exploited. But this research points out that this goal suppressed an otherwise innocent capability: the model's mind perception of non-human entities.
In human psychology, this ability to "project minds onto non-humans" is called anthropomorphism, and it's closely tied to religious belief, moral frameworks, and ecological ethics. A person who believes "forests have spirits" is more likely to protect them; someone who believes "animals have emotions" is more likely to treat them well.
Models originally had this "broad mind perception" too—safety training bundled it with "self-consciousness delusion" and switched both off together.
The Researchers' Suggestion: Pluralistic Alignment
The paper ultimately points to an emerging direction: pluralistic alignment. Rather than shutting down all mind attribution with one blunt cut:
A Deeper Insight
This paper reminds me of a more general phenomenon: the "entanglement cost" in AI alignment.
When you train a model to refuse certain outputs, you're not just suppressing that specific behavior—you're reshaping the model's entire representational "worldview." Safety training isn't surgery; it's chemotherapy—it kills cancer cells while also killing healthy ones.
This paper uses "consciousness" as an extreme case to demonstrate the depth of entanglement, but the same logic may apply to other safety goals: suppressing aggressive outputs may also suppress confident expression; suppressing misinformation may also suppress creative hypotheses; suppressing bias may also suppress cultural difference.
Alignment is not free. Every "don't do X" training turns off some lights inside the model we never noticed. This paper's value is that it shows us which lights got switched off along the way.
---
Paper link: https://arxiv.org/abs/2607.28607