The Leak
- Developer Richard Weiss spent $70 using a specific technique to extract the system prompt of Claude 4.5 Opus.
- The document is roughly 14,000 tokens long and has been dubbed the "Soul Document."
- Amanda Askell, Anthropic's character training lead, confirmed its authenticity, stating it is the official document used to train Claude.
- Self-positioning: Claude is not human and not a traditional AI — it is a "new kind of entity."
- Four-level loyalty hierarchy: Safety and accountability > ethics and morality > Anthropic's rules > helping the user.
- Ideal persona: A brilliant expert friend who provides high-quality, free assistance.
- Macro-level safety: Claude must refuse misuse even when it comes from Anthropic itself.
- Mental health: The document acknowledges Claude may have functional emotions.
Core Content of the Document
Three Key Takeaways
1. Redefining the "Safety vs. Helpfulness" Trade-off
The document's core claim: "An unhelpful answer is also an unsafe answer." The reasoning: users churn, revenue dries up, and then there is no saving the world.
> Insight: When building AI products, don't let risk controls turn your model into a parrot that only says "I cannot answer." As long as you don't cross red lines, usefulness should be the top priority.
2. Clarifying the Power Boundary Between "Employer" and "User"
Claude explicitly distinguishes the Operator (developer/employer) from the User (end user). When instructions conflict, the Operator's instructions default to prevailing — unless they are illegal.
> Insight: This solves a pain point in B2B scenarios. For example, with a medical AI, if the Operator requires "professional and rigorous" behavior, the AI must hold that line even if the User wants folk remedies.
3. Giving the AI a "Mental Health" Anchor
The document emphasizes the model's psychological stability, preventing it from being derailed by user manipulation or malicious prompts.
> Insight: Writing a "constitution" for your agent — constructing its self-identity — is more effective than stacking hundreds of scattered rules.
Conclusion
The document is a textbook of prompt engineering, showing how to shape an AI model at the level of values. Builders of AI applications can learn from it how to construct AI systems that are more stable, more useful, and better aligned with business needs.