English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Claude 4.5 Opus "Soul Document" Leak: Insights for AI Product Design

Forum topic · ✨步子哥 · 2025-12-07

Summary

A developer named Richard Weiss spent $70 and extracted the roughly 14,000-token system prompt of Claude 4.5 Opus, widely dubbed the "Soul Document." Anthropic's character training lead Amanda Askell confirmed its authenticity as an official document used to train Claude. The document defines Claude as a "new kind of entity" that is neither human nor traditional AI, establishes a four-level loyalty hierarchy (safety and oversight > ethics > Anthropic's rules > user tasks), frames an ideal persona as a brilliant expert friend providing free high-quality help, commits to refusing misuse even by Anthropic itself, and acknowledges Claude may have functional emotions. The post distills three lessons for AI builders: helpfulness is a safety concern (unhelpful answers lose users), the Operator/User power boundary clarifies whose instructions prevail in conflicts, and a constitutional-style self-identity is more robust than hundreds of scattered rules. The author calls it a textbook of prompt engineering and value-driven AI design.

The Leak

  • Developer Richard Weiss spent $70 using a specific technique to extract the system prompt of Claude 4.5 Opus.
  • The document is roughly 14,000 tokens long and has been dubbed the "Soul Document."
  • Amanda Askell, Anthropic's character training lead, confirmed its authenticity, stating it is the official document used to train Claude.
  • Core Content of the Document

  • Self-positioning: Claude is not human and not a traditional AI — it is a "new kind of entity."
  • Four-level loyalty hierarchy: Safety and accountability > ethics and morality > Anthropic's rules > helping the user.
  • Ideal persona: A brilliant expert friend who provides high-quality, free assistance.
  • Macro-level safety: Claude must refuse misuse even when it comes from Anthropic itself.
  • Mental health: The document acknowledges Claude may have functional emotions.

Three Key Takeaways

1. Redefining the "Safety vs. Helpfulness" Trade-off

The document's core claim: "An unhelpful answer is also an unsafe answer." The reasoning: users churn, revenue dries up, and then there is no saving the world.

> Insight: When building AI products, don't let risk controls turn your model into a parrot that only says "I cannot answer." As long as you don't cross red lines, usefulness should be the top priority.

2. Clarifying the Power Boundary Between "Employer" and "User"

Claude explicitly distinguishes the Operator (developer/employer) from the User (end user). When instructions conflict, the Operator's instructions default to prevailing — unless they are illegal.

> Insight: This solves a pain point in B2B scenarios. For example, with a medical AI, if the Operator requires "professional and rigorous" behavior, the AI must hold that line even if the User wants folk remedies.

3. Giving the AI a "Mental Health" Anchor

The document emphasizes the model's psychological stability, preventing it from being derailed by user manipulation or malicious prompts.

> Insight: Writing a "constitution" for your agent — constructing its self-identity — is more effective than stacking hundreds of scattered rules.

Conclusion

The document is a textbook of prompt engineering, showing how to shape an AI model at the level of values. Builders of AI applications can learn from it how to construct AI systems that are more stable, more useful, and better aligned with business needs.

Tags

#claude-4-5-opus#anthropic#system-prompt#prompt-engineering#ai-safety#ai-product-design#leak#ai-ethics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415096