HEART-Bench: Giving AI a "Psychological Exam" — 11 Virtual Personas, 1000 Memories, 673 Multiple-Choice Questions
> Source: HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?, Weihan Peng, Chenxu Zhang, Qianao Wang, Yuling Shi, Heng Lian et al., https://arxiv.org/abs/2605.30058
---
Have you ever wondered: if you ask an LLM to play a "high-neuroticism, low-extraversion" person, can it stay consistent across all scenarios?
Not the one-dimensional acting where it just adds exclamation marks to every sentence after being told "play an anxious person" — but giving it a full life story and then checking, across workplace, romance, social, and crisis scenarios, whether its decisions actually follow that person's psychological logic.
That is exactly what HEART-Bench measures.
Not "Role-Playing" but "Becoming"
Existing LLM personality tests are mostly simplistic: give a prompt describing a personality, ask a few questions, and check if the answers match. The problem is that this only tests "acting ability," not "personality consistency" — a model might act introverted in the first question and forget who it is by the fifth.
HEART-Bench takes a completely different approach. It builds 11 complete human personas, each with:
1. Orthogonal Big Five traits: openness, conscientiousness, extraversion, agreeableness, and neuroticism vary independently across personas, ensuring fundamental differences between them. 2. 1000 autobiographical episodic memories: not dry personality descriptions, but concrete, emotionally warm memories like "the first time I spoke in front of the whole school at age seven, my palms were sweating from nervousness." These memories are distributed across life stages — childhood, adolescence, early adulthood, middle age — based on developmental psychology theory. 3. Human validation: all persona settings and questions are manually screened to ensure psychological plausibility.
The DIAMONDS Framework: Scanning Personality Across 8 Dimensions
With personas in place, you need test scenarios. HEART-Bench uses the DIAMONDS taxonomy from psychology, dividing situations into 8 dimensions:
| Dimension | Meaning | Example scenario | |------|------|---------| | Duty | Responsibility and obligation | Completing work tasks on time | | Intellect | Intellectual challenge | Thinking deeply about complex problems | | Adversity | Adversity and stress | Coping after failure | | Mating | Romance and intimacy | Reacting to a confession of love | | pOsitivity | Positive situations | Responding to an unexpected gift | | Negativity | Negative situations | Handling being misunderstood | | Deception | Deception and trust | Discovering a friend lied | | Sociality | Social interaction | Behaving at a party |
Each persona must make behavioral decisions across all 8 categories, producing 673 multiple-choice questions, each human-verified.
What Is Being Measured? Not Knowledge, but "Personality Consistency"
The core metric of HEART-Bench is not "how many questions are answered correctly," but: does a high-neuroticism person lean toward anxious reactions in adversity more than a low-neuroticism one? Does a high-extraversion person act more proactively in social settings than a low-extraversion one?
This mirrors "construct validity" in psychology experiments — looking not at surface scores but at whether behavioral patterns match theoretical expectations.
Why Does This Matter?
Current AI agent research focuses almost entirely on "task capability" — reasoning, planning, tool use. But a genuinely useful AI assistant needs not only to complete tasks; it also needs stable personality traits. You wouldn't want your assistant to be warm and considerate today and cold and indifferent tomorrow.
The deeper implication: if an LLM cannot maintain a consistent personality, it is unreliable in any scenario requiring long-term trust — psychological counseling, educational tutoring, customer service. The core of these scenarios is not "can it do the job" but "can people trust it."
HEART-Bench sends a clear signal: we need not only smarter AI, but psychologically more "complete" AI. Measuring that completeness requires experiment design as rigorous as psychology — not casually asking a few questions, but building full personas, designing systematic scenarios, and verifying behavioral consistency.
673 questions is not many, but each one is backed by personality theory, developmental psychology, and human validation. This is bringing the methodology of psychology experiments into AI evaluation — not making AI take a test, but giving AI a psychological exam.
---
Paper link: https://arxiv.org/abs/2605.30058