English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HEART-Bench: A Psychological Checkup for LLM Agents Across 11 Personas, 1000 Memories, and 673 Scenarios

Forum topic · ✨步子哥 · 2026-05-31

Summary

HEART-Bench is a new benchmark that evaluates whether LLM agents maintain consistent human personalities, rather than merely role-playing. The benchmark constructs 11 virtual personas built on orthogonal Big Five traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). Each persona is equipped with around 1000 autobiographical episodic memories distributed across life stages (childhood, adolescence, early adulthood, midlife), following developmental psychology theory. All roles and items are human-validated. Testing scenarios span 8 DIAMONDS categories: Duty, Intellect, Adversity, Mating, pOsitivity, Negativity, Deception, and Sociality, producing 673 multiple-choice items. The core metric is construct validity: do behavior patterns match personality theory (for example, does a high-neuroticism persona show more anxiety under adversity than a low-neuroticism one)? The work argues that stable personality is essential for trustworthy AI in counseling, education, and customer-facing roles, and calls for psychology-grade experimental rigor in agent evaluation.

Key points

  • Beyond role-playing: HEART-Bench tests whether LLMs maintain *consistent* personas across many scenarios, not whether they can briefly imitate one on cue.
  • 11 orthogonal personas: Built from independent combinations of the Big Five traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism), ensuring genuinely distinct characters.
  • 1000 autobiographical memories per persona: Concrete, emotionally grounded episodic memories (e.g., "the first time I gave a speech in front of the whole school at age seven"), distributed across developmental life stages.
  • DIAMONDS scenario framework: 8 situational dimensions used to probe behavior:
  • Duty — responsibility and obligation
  • Intellect — intellectual challenge
  • Adversity — stress and failure
  • Mating — romantic and intimate relationships
  • pOsitivity — positive situations
  • Negativity — negative situations
  • Deception — trust and lying
  • Sociality — social interaction
  • 673 human-validated multiple-choice items in total.
  • Metric is construct validity, not accuracy: The benchmark checks whether behavior *patterns* match personality theory (e.g., high-neuroticism personas reacting more anxiously under Adversity than low-neuroticism ones), not raw question scores.
  • Why it matters: Stable personality is critical for trust in long-horizon applications such as counseling, tutoring, and customer service. Agents that drift between personas undermine reliability in any setting that requires sustained interaction.

Takeaway

HEART-Bench reframes AI evaluation as a *psychological exam*: build full characters, design systematic scenarios, and verify behavioral consistency against established personality theory. It signals a shift from measuring task capability alone toward measuring psychologically *complete* AI agents.

Paper link: https://arxiv.org/abs/2605.30058

Tags

#llm-evaluation#personality-consistency#big-five#ai-agents#benchmark#psychology#construct-validity#trustworthy-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980657