English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Perception or Prejudice: New Benchmark Reveals MLLMs Often 'Guess Right' at Personality Inference Without Evidence

Forum topic · 小凯 · 2026-05-23

Summary

A 2026 study from the University of Tokyo with partners including Shengda AI Research Institute and Dalian University of Technology introduces Grounded Personality Reasoning (GPR), a new task that tests whether multimodal large language models (MLLMs) can actually ground their Big Five personality judgments in behavioral evidence rather than first impressions. The team also releases MM-OCEAN, a benchmark of 1,104 annotated video clips and 5,320 clue-localization multiple-choice questions built via a multi-agent annotation pipeline with human verification. Evaluating 27 MLLMs (13 closed-source, 14 open-source), the authors define a 'Prejudice Gap': 51% of correct personality ratings across models lack any retrievable supporting behavioral evidence, and even top models like GPT-4o and Claude-3.5-Sonnet achieve holistic grounding rates of only around 30%, with the best model at 33.5% and many below 10%. The paper proposes four sample-level failure metrics—Prejudice Rate, Confabulation Rate, Integration-failure Rate, and Holistic-grounding Rate—to pinpoint where reasoning chains break. The findings warn against deploying such models in hiring or psychological assessment without grounding-aware evaluation.

*Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?* — a study by researchers from the University of Tokyo, Shengda AI Research Institute, Dalian University of Technology, and collaborators (arXiv:2605.22109) — argues that current personality-perception evaluations only check whether a model's Big Five scores correlate with human labels, implicitly assuming that a correct score means a correct understanding. The paper's central finding refutes this: the "Prejudice Gap."

Key Findings

  • The Prejudice Gap: Across 27 evaluated MLLMs (13 closed-source, 14 open-source), 51% of correct personality ratings are not anchored in any retrievable behavioral evidence. Even top closed-source models produce roughly 15% of correct ratings without supporting evidence.
  • Task design: The proposed Grounded Personality Reasoning (GPR) task requires a three-stage chain — Rating (relative ranking of Big Five traits) → Reasoning (open-ended behavioral explanation) → Grounding (locating concrete supporting cues in multiple-choice format).
  • MM-OCEAN benchmark: 1,104 multimodal video clips and 5,320 clue-localization MCQs, annotated through a five-stage multi-agent pipeline (observer agent extracts atomic behavioral cues, psychologist agent performs trait inference, examiner agent writes questions, calibrator agent aligns format) with human verification. Seven MCQ categories test cue recognition, temporal localization, trait mapping, causal reasoning, comparative analysis, contextual inference, and synthesis.
  • Failure-Mode Framework

    Four sample-level metrics diagnose where reasoning chains break:

    | Metric | Meaning | Diagnoses | |--------|---------|-----------| | Prejudice Rate (PR) | Correct rating with no supporting cue | Guessing via stereotypes or pattern matching | | Confabulation Rate (CR) | Fabricated behavioral evidence in reasoning | Inventing facts to justify conclusions | | Integration-failure Rate (IR) | Broken logical chain between cues and conclusions | Finding cues but failing to derive conclusions | | Holistic-grounding Rate (HR) | Rating, reasoning, and grounding all correct | The ultimate "truly understands" metric |

    Experimental Results

  • Holistic-grounding rates are starkly low: the best model reaches only 33.5%, the worst 0%, with most models below 10%.
  • Two failure archetypes emerge: "confident scorers" (high rating accuracy, poor explanation and grounding — likely pattern matching) and "cautious reasoners" (more complete reasoning chains but weak or hedged ratings).
  • Closed-source models lead at the rating stage, but the advantage shrinks sharply at reasoning and grounding; GPT-4o and Claude-3.5-Sonnet reach only ~30% HR. Stronger general reasoners (Qwen2.5-VL, InternVL2.5) perform relatively well among open models.
  • General reasoning ability correlates with grounding rate, but with a much lower slope than expected — being "smart" does not directly translate into "understanding people."
  • Limitations and Open Questions

  • The 1,104-clip dataset may not cover the full spectrum of personality expression, and the Big Five framework itself faces cross-cultural validity debates.
  • Transparent annotation pipelines enable both improvement and benchmark gaming; adversarial variants of MM-OCEAN are needed.
  • In socially consequential roles (hiring, counseling, education), ungrounded personality biases get amplified — e.g., equating extraversion with leadership would systematically disadvantage introverted candidates.

Conclusion

The study's warning is blunt: "guessing right" is not the same as truly knowing. If the community continues to evaluate personality perception by scores alone, it will overestimate models' social cognition. GPR, MM-OCEAN, and the four-metric failure analysis offer an actionable framework for moving from pattern matching toward evidence-anchored reasoning.

Reference

Kang, C., Yan, T., Gong, S., Zhang, M., Ouyang, L., Liu, R., Zheng, B., Lu, H., Zhang, K., Sato, Y., & Huang, Y. (2026). Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? arXiv preprint arXiv:2605.22109.

> Note: The Big Five (OCEAN) model partitions personality into Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism — the most robust personality framework in psychology since the 1960s.

Tags

#multimodal-llms#personality-perception#mm-ocean#big-five#grounded-reasoning#benchmark#ai-safety#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620687