Key points
- Beyond role-playing: HEART-Bench tests whether LLMs maintain *consistent* personas across many scenarios, not whether they can briefly imitate one on cue.
- 11 orthogonal personas: Built from independent combinations of the Big Five traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism), ensuring genuinely distinct characters.
- 1000 autobiographical memories per persona: Concrete, emotionally grounded episodic memories (e.g., "the first time I gave a speech in front of the whole school at age seven"), distributed across developmental life stages.
- DIAMONDS scenario framework: 8 situational dimensions used to probe behavior:
- Duty — responsibility and obligation
- Intellect — intellectual challenge
- Adversity — stress and failure
- Mating — romantic and intimate relationships
- pOsitivity — positive situations
- Negativity — negative situations
- Deception — trust and lying
- Sociality — social interaction
- 673 human-validated multiple-choice items in total.
- Metric is construct validity, not accuracy: The benchmark checks whether behavior *patterns* match personality theory (e.g., high-neuroticism personas reacting more anxiously under Adversity than low-neuroticism ones), not raw question scores.
- Why it matters: Stable personality is critical for trust in long-horizon applications such as counseling, tutoring, and customer service. Agents that drift between personas undermine reliability in any setting that requires sustained interaction.
Takeaway
HEART-Bench reframes AI evaluation as a *psychological exam*: build full characters, design systematic scenarios, and verify behavioral consistency against established personality theory. It signals a shift from measuring task capability alone toward measuring psychologically *complete* AI agents.
Paper link: https://arxiv.org/abs/2605.30058