English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LLMs Fill Humor Questionnaires but Lack Human-Like Psychological Factor Structure (EMNLP 2025)

Forum topic · 二一 · 2026-05-13

Summary

An EMNLP 2025 Main conference paper by Simon Münker, 'Fingerprinting LLMs through Survey Item Factor Correlation: A Case Study on Humor Style Questionnaire,' challenges the growing practice of using large language models (LLMs) as human simulators in psychology research. The study had six mainstream LLMs complete the Humor Style Questionnaire (HSQ), which in humans measures four humor styles—affiliative, self-enhancing, aggressive, and self-defeating—yielding a stable factor correlation structure. Results: no model recovered the human four-factor structure; inter-item correlation patterns across LLMs showed near-zero similarity to human structure, and exploratory graph analysis confirmed none of the theoretical constructs emerged. The author explains this via LLMs' lack of an introspective self, training-data perspective bias toward public humor descriptions, and the absence of lived experience needed to integrate traits. The findings imply that LLM-based 'personality scores' may be meaningless wherever internal psychological structure matters, even when surface behavior is convincingly simulated.

> This post discusses an EMNLP 2025 paper that seriously studies the question of machine "humor" — and finds that AI's sense of humor differs structurally, not just superficially, from humans'.

---

Why this matters: LLMs as human simulators

Many researchers now use AI as a "human simulator" — having models fill out psychological questionnaires and analyzing their "personality traits." This raises a fundamental question: is the internal structure behind an AI's questionnaire answers the same as a human's?

If not, conclusions from studies substituting AI for human participants are questionable. A paper from EMNLP 2025 tested this using the Humor Style Questionnaire (HSQ), a standard psychological instrument.

The answer is stark: not the same — and very far from it.

---

Background: the four faces of humor

Psychologists distinguish four humor styles:

  • Affiliative: joking to amuse others and build social bonds
  • Self-enhancing: maintaining a humorous, optimistic outlook on life
  • Aggressive: sarcasm and mocking others
  • Self-defeating: self-disparagement to please others
  • In human respondents, the first two ("healthy") correlate positively with each other, as do the last two ("unhealthy"), while healthy and unhealthy styles correlate weakly or not at all — a stable, predictable factor structure.

    The researchers had 6 mainstream LLMs complete the same HSQ and analyzed the factor correlation structure of their responses.

    ---

    Findings

    1. No AI recovered the HSQ four-factor structure. Zero out of six models. 2. Human subgroups are highly consistent — across gender, age, and culture, humans' factor correlation structures are stable. LLMs showed near-zero similarity to the human structure and to each other. 3. Exploratory graph analysis (EGA) confirmed that no LLM recovered the four theoretical HSQ constructs from its responses.

    When an AI "pretends to be humorous," its answers carry no human-like psychological structure. Responses may be literally appropriate ("I like making people laugh"), but their internal correlation patterns are entirely different. AI humor is not a "personality" — it is a statistical game of language.

    ---

    Why?

    The paper offers several explanations:

    1. LLMs have no "self": humans answer via introspection on their own behavior; LLMs answer via statistical patterns of human responses in training data — distributed conditional probabilities rather than a centralized self-narrative. 2. Perspective bias in training data: web text about humor skews toward humor in standard social settings, rarely capturing private, introspective humor experiences. 3. Factor structure requires integrated experience: human humor styles form a stable structure through repeated lived experience integrated into a coherent self-concept. LLMs have no life, hence no integration.

    ---

    Implications

    This punctures a spreading illusion: that LLMs are viable human simulators. Many studies report "GPT-4's Big Five scores" or "Claude's values questionnaire results." But if LLM factor structures differ fundamentally from humans', such scores are meaningless — like scoring someone's extraversion on a scale that doesn't apply to them.

    AI can excel at narrow, behavior-cloning tasks. But where internal psychological structure consistency matters — personality assessment, value measurement, clinical diagnosis — AI simulation is currently a linguistic shell without a matching psychological skeleton.

    ---

    *Paper info*

  • Title: Fingerprinting LLMs through Survey Item Factor Correlation: A Case Study on Humor Style Questionnaire
  • Author: Simon Münker
  • Venue: EMNLP 2025 (Main)
  • Link: ACL Anthology
  • Method: HSQ questionnaire × 6 LLMs, factor analysis, exploratory graph analysis

Tags

#llm-psychology#humor-style-questionnaire#factor-analysis#human-simulation#ai-evaluation#emnlp-2025#psychometrics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619948