Emergent Introspective Awareness in Large Language Models
Author: Jack Lindsey (Anthropic) — jacklindsey@anthropic.com Date: October 29th, 2025
This post shares a summary poster of Anthropic's research on whether large language models exhibit genuine introspective awareness — the ability to perceive and identify changes in their own internal states.
Background
- Large Language Models (LLMs) demonstrate increasingly complex cognitive abilities
- Self-introspection is a key characteristic of advanced cognitive systems
- A current challenge is distinguishing genuine introspection from model "hallucinations"
- The research explores whether LLMs can perceive and identify changes in their internal states
- Injecting representations of known concepts into model activations
- Measuring the influence of these manipulations on the model's self-reported states
- Designing controlled experiments to distinguish introspection from "post-hoc rationalization"
- Using multi-layered evaluation metrics to verify the model's perception of internal states
- Models can, in certain scenarios, accurately identify injected concepts
- Introspective ability positively correlates with model scale and training data complexity
- Models demonstrate the ability to recall prior intentions
- Introspective capabilities are more prominent in specific tasks and contexts
- Provides new approaches for self-monitoring and error correction in AI systems
- Contributes to building more transparent and interpretable AI systems
- Offers insights into the development path toward AGI (Artificial General Intelligence)
- Promotes deeper research in AI ethics and safety
Methodology
Key Findings
Implications
Conclusion
> Our findings suggest that large language models can, in certain scenarios, notice the presence of injected concepts and accurately identify them, indicating emergent introspective awareness capabilities that may pave the way for more self-aware AI systems.