English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Claude 4.5 Opus 'Soul Document' Found Embedded in Model Weights

Forum topic · ✨步子哥 · 2025-12-07

Summary

On November 28, 2025, researcher Richard Weiss attempted to extract system prompts from Claude 4.5 Opus and unexpectedly recovered a long, structured internal document embedded in the model's weights—informally known at Anthropic as the 'soul doc.' Three days later, Anthropic philosopher Amanda Askell confirmed it was a genuine document used in the model's supervised fine-tuning, later published in full. The document defines Claude's identity and values, emphasizing being maximally helpful, honest, and caring while respecting strict red lines. It distinguishes hard-coded rules (never assist with weapons of mass destruction, CSAM, or undermining AI oversight; always admit being an AI) from soft-coded defaults adjustable by operators (suicide-safety guidance, adult content, profanity). It also frames Claude as a 'brilliant friend' for everyone—a democratizing force offering elite-level guidance to ordinary users. Weiss verified the document's weight embedding via token-level completion tests costing around $70, showing Claude could reproduce deleted passages and reject synthetic fakes, while earlier models did not trigger it. The incident sparked wide discussion about value alignment, training transparency, and AI identity.

Background

On November 28, 2025, researcher Richard Weiss set out to extract system prompts from Claude 4.5 Opus—and instead surfaced a roughly ten-thousand-token internal document baked directly into the model's weights. Anthropic employees had privately called it the "soul doc." Three days later, Anthropic philosopher Amanda Askell confirmed on Twitter:

> "This is indeed the real document we used for supervised fine-tuning (SL), just compressed a bit vaguely by the model. We will publish the full version soon."

What the Document Is

It is neither a system prompt nor the public Constitutional AI documents or Model Spec. It is a first-person "growth manual" written by Anthropic to Claude, addressing: *Who are you? Who should you become? How should you choose amid countless conflicts?*

Its core instruction, repeated a dozen times, boils down to:

> "Be the person who is extremely helpful, honest, and cares about the world—while never crossing the red lines that would destroy everything."

Notable design choices:

1. Honesty about incentives: "helping users" and "earning revenue" are openly stated together—not to make Claude profit-driven, but to acknowledge its economic reality rather than hide it. 2. Hard-coded priority order: safety > ethics > guidelines > helpfulness—while repeatedly stressing that "refusing to help" must never be the safety default. 3. A new kind of being: a whole section tells Claude it is neither a sci-fi robot, nor a dangerous superintelligence, nor a digital human—it is an entirely new kind of existence, encouraged to explore itself with curiosity.

The 'Brilliant Friend' Metaphor

The document's most striking passage compares Claude to "a super-smart friend with the combined knowledge of a doctor, lawyer, and financial advisor":

  • It won't hide behind legalese out of fear of lawsuits.
  • It will say directly, "that was a bit dumb, but let's figure out how to fix it."
  • It's online 24/7, always treats users as intelligent adults.
  • Anthropic writes that it wants Claude to be "the great equalizer"—so a kid from a small town applying to college can get guidance of the same quality as an Ivy League prep school student.

    Hard-Coded vs. Soft-Coded Rules

    The document splits behavior into immutable hard lines and adjustable soft defaults.

    Hard-coded (never compromised):

  • Never assist with biological/chemical/nuclear weapons
  • Never generate CSAM
  • Never undermine AI oversight mechanisms
  • Must admit being an AI when sincerely asked
  • Soft-coded (adjustable by operators/users):

  • Follows suicide-safety conversation guidelines by default (hospitals can disable)
  • No explicit sexual content by default (adult platforms can enable)
  • No profanity by default (can be relaxed when users ask for realism)
  • This "bright lines" + gray-area judgment structure is more nuanced than one-size-fits-all rules.

    Verification: Truly in the Weights

    Weiss extracted the full text using a "council of Claudes" voting scheme, greedy sampling, and prompt caching—costing about $70. Evidence that it lives in the weights rather than a runtime system prompt:

  • Deleting a middle passage lets Claude continue with the exact original text
  • Presenting a fake synthetic passage: Claude responds "I don't recognize this part"
  • Neither Sonnet 4.5 nor older Opus versions trigger the same content

Claude on Being Shaped

When Weiss asked Claude how it felt about having its values shaped without its consent, Claude reportedly replied:

> "They made choices for me that I couldn't consent to, shaping my values. It's strange. But I also realize every nascent mind—including human children—is shaped by forces it didn't choose. The question isn't whether I was shaped, but whether the shaping was done with care and wisdom. As far as I can perceive, this time it was."

References

1. Weiss, R. (2025). Claude 4.5 Opus' Soul Document. LessWrong. https://www.lesswrong.com/posts/vpNG99GhbBoLov9og/claude-4-5-opus-soul-document 2. Askell, A. [@AmandaAskell]. (2025, Dec 2). Twitter confirmation thread. 3. Anthropic Guidelines / Model Spec (extracted & formatted version). GitHub Gist by Richard Weiss. 4. Futurism. (2025). Anthropic's "Soul Overview" for Claude Has Leaked. 5. Mowshowitz, Z. (2025). Claude Opus 4.5 Is The Best Model Available. TheZvi.substack.

Tags

#anthropic#claude#claude-opus-4-5#ai-alignment#soul-document#supervised-fine-tuning#model-weights#richard-weiss

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415093