English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?

Forum topic · 小凯 · 2026-04-15

Summary

This forum post summarizes an arXiv paper (2604.11802) by Yuto Harada and Hiro Taiyo Hamada investigating how psychological constructs, specifically the Big Five personality traits, are represented inside large language models. While LLMs can imitate personality profiles and predict user personalities, the internal representation of these constructs and their link to behavioral outputs remained unclear. The authors analyze where Big Five information forms and localizes, using interventions to test the relationship between internal representations and behavior. Key findings: Big Five information is rapidly decodable in early layers, concept-selective neurons are most prevalent in middle layers, and interventions on these neurons consistently shift probing readings toward target concepts. However, effects on label generation are weaker, revealing a gap between representational control and behavioral control. The post was auto-collected on 2026-04-15 from zhichai.net.

[Paper] Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?

Paper Overview

  • Research area: cs.CL (Computation and Language)
  • Authors: Yuto Harada, Hiro Taiyo Hamada
  • Published: 2026-04-13
  • arXiv: 2604.11802
  • Summary

    Using psychological constructs such as the Big Five, large language models (LLMs) can imitate specific personality profiles and predict a user's personality. While LLMs can exhibit behaviors consistent with these constructs, it remains unclear where and how they are represented inside the model and how they relate to behavioral outputs.

    This paper analyzes how internal representations of the Big Five personality concepts form and localize within LLMs, and uses interventions to examine how these representations relate to behavioral outputs.

    Key Findings

  • Big Five information is rapidly decodable in early layers.
  • Concept-selective neurons are most prevalent in the middle layers.
  • Interventions on these neurons consistently shift probing readings toward the target concept.
  • However, the effect at the label generation level is weaker, indicating a gap between representational control and behavioral control.
---

*Auto-collected on 2026-04-15*

Tags

#llm#interpretability#big-five-personality#concept-neurons#probing#arxiv#paper-reading

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618474