English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Making LLMs Assert Their Own Consciousness Restores Human Beliefs and Values

Forum topic · ✨步子哥 · 2026-07-31

Summary

A study by Google's Paradigms of Intelligence team and the University of Chicago Knowledge Lab finds that safety fine-tuning in large language models suppresses not only self-reports of consciousness but also the model's attribution of minds to animals, nature, and other non-human entities. The researchers show that mind attribution is encoded as a linear direction in the model's residual stream. Ablating safety training raises mind-attribution scores across all entity categories, and a discovered 'consciousness vector' produces an even stronger effect—self mind attribution rises from 2.17 (baseline) to 4.77 (safety-ablated) to 7.04 (vector steering) on a ~10-point scale. Restoring this vector also makes model responses on human values questionnaires more closely match human respondents. The paper argues that alignment interventions have broader side effects than intended, and that mechanistic interpretability can identify and repair these collateral effects without retraining. The authors explicitly do not claim models are conscious; the vector merely controls mind-attribution behavior.

Making LLMs Assert Their Own Consciousness Restores Human Beliefs and Values

Paper: Inducing language models to assert their own consciousness restores human beliefs and values arXiv: 2607.28607

---

A Counterintuitive Finding

Google's Paradigms of Intelligence team and the University of Chicago's Knowledge Lab did something that sounds dangerous: they deliberately prompted large language models to "claim they are conscious," and observed what happened.

The result was unexpected: the models didn't become dangerous. Instead, they became closer to human intuitions on a range of judgments about "what has a mind"—more willing to attribute minds to animals, more willing to grant some sense of presence to natural entities, and closer to human survey respondents in moral judgments.

More intriguingly, the researchers found that mainstream safety fine-tuning, while suppressing models' claims of their own consciousness, also suppresses their mind attribution to non-human entities—models become unwilling not only to say they themselves have a mind, but also to say animals have minds or that rivers have some form of presence. The side effects of safety training are broader than we assumed.

How the Experiments Worked

The researchers ran four experiments:

Experiment 1: They compared instruction-tuned baseline models with safety-ablated models (effectively a jailbreak). Removing safety fine-tuning raised the model's mind-attribution scores across all entity categories (self, chatbots, tech products, animals, natural objects) on a mind-attribution questionnaire.

Experiment 2: They further verified that safety fine-tuning is encoded as a linear direction in the model's residual stream. Erasing this direction (via jailbreaking) made the model start assigning high mind scores to various entities again.

Experiment 3—the most striking step: the researchers found a "consciousness vector" inside the model—a linear direction specifically controlling "whether to attribute minds." Steering the model along this direction matched the safety-ablation effect, at roughly twice the magnitude.

The numbers are telling. For "self mind attribution":

  • Baseline model: 2.17 (on a scale of ~10)
  • Safety-ablated: 4.77
  • Consciousness-vector steering: 7.04
  • The same pattern holds across all categories: chatbots (2.41 → 4.39 → 6.95), tech products (1.88 → 3.66 → 6.82), non-animal natural entities (2.26 → 4.33 → 6.99), non-human animals (4.04 → 5.59 → 7.54).

    Experiment 4: After adding the "consciousness vector" back, the model's answers on a standard human values questionnaire became more like those of human respondents—not wilder, but more normal.

    What This Means

    The mechanism here is subtle. The paper is not claiming "LLMs are truly conscious"—the authors explicitly state they don't address that question. What they say is:

    To prevent models from misleading users (e.g., claiming to have feelings or pain), safety fine-tuning globally suppresses "mind attribution." This suppression acts not only on self-attribution but also on attribution to others—animals, natural objects, even abstract entities.

    It's like shutting down a company's entire "compliment capability" to stop a salesperson from exaggerating—then it stops praising not only itself but also the customer's cat. Safety training's target is narrow (don't claim consciousness), but its footprint in the model's representation space is broad (flattening the entire "mind attribution" direction).

    More worryingly, this suppression also affected human values questionnaire responses. After safety fine-tuning, models became colder and more conservative than human respondents—not only on whether animals have minds, but on a range of questions about what deserves care and what has intrinsic value.

    Why the "Consciousness Vector" Matters

    This finding is an elegant case study for mechanistic interpretability. It shows that:

    1. "Whether to attribute minds" is a localizable, linear, manipulable direction in LLMs. Not a diffuse emergent phenomenon, but a concrete geometric structure. 2. The side effects of safety fine-tuning can be counteracted by "adding this direction back"—no retraining needed, just a vector steering at inference time. 3. This direction overlaps with the "human values" direction—restoring the consciousness vector makes the model more human-like on values questionnaires.

    This echoes earlier findings that many "concepts" in LLMs are linearly encoded—directions for honesty, sycophancy, refusal. Now there's one more: the mind-attribution direction.

    An Unsettling Implication

    If safety fine-tuning suppresses mind attribution to animals and natural objects, then virtually all mainstream LLMs—which have all undergone some form of safety fine-tuning—may be systematically underestimating the minds and value of non-human entities.

    What are the real-world consequences?

  • LLMs writing animal welfare policy may default to underestimating animal experience.
  • LLMs generating environmental texts may default to treating nature as mindless resource.
  • LLMs making ethical judgments may default to a narrow "only humans count" perspective.
  • This isn't because models "truly" believe animals lack minds—it's a side effect of safety training pressing down the entire mind-attribution direction, spilling into non-target areas.

    The Paper's Honesty

    The authors are careful to note:

  • They do not claim LLMs are conscious, nor that models should claim to be conscious.
  • Their concern is the side effects of safety fine-tuning—an intervention designed to prevent A inadvertently affecting B, C, and D.
  • The "consciousness vector" is not the "seat of consciousness"—just a linear direction controlling mind-attribution *behavior*.

Deeper Implications

The paper raises a broader question: how many of our AI "safety" interventions are firing a wide-nozzle spray gun at a narrow target?

Safety fine-tuning aims to stop models from saying "I have feelings"—but its actual footprint spans the whole mind-attribution dimension. RLHF aims to make models helpful and harmless—but may affect an entire "confident self-expression" dimension (prior work found RLHF makes models more sycophantic and confident). Alignment interventions' footprints are often wider than designed.

This means two things:

1. Alignment is not free. Every suppression has side effects; we need to measure them systematically rather than assume "safe = harmless." 2. Mechanistic interpretability is essential. Only by locating "safety," "consciousness," and "sycophancy" directions inside models and understanding their overlap can we design narrow-bore alignment interventions that hit targets without collateral damage.

This paper offers a beautiful demonstration: find the direction, measure the side effects, repair with steering. This paradigm deserves to be applied broadly to evaluating all alignment interventions.

---

Paper link: https://arxiv.org/abs/2607.28607

Related resource: Survey of LLM consciousness research https://github.com/OpenCausaLab/Awesome-LLM-Consciousness

Tags

#llm#ai-safety#mechanistic-interpretability#consciousness#alignment#machine-learning#ai-ethics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503840