English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind - Paper Review

Forum topic · 小凯 · 2026-05-21

Summary

This zhichai.net forum post reviews the paper OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind (arXiv:2505.10250), which targets a major weakness of large language models: reasoning about nested beliefs and information asymmetry, known as high-order Theory of Mind (ToM). The authors introduce an Observer-Self Conflict setup where an agent must both know a fact and hide that knowledge, a scenario that cripples prior systems. OSCToM uses a reinforcement-learning-guided adversarial generator built on an extended domain-specific language to produce ToM puzzles that the target model fails on, then trains an 8B-parameter model with compositional surrogate models that split ToM into interpretable sub-skills. On the FANToM benchmark, accuracy jumps from ExploreToM's 0.2% to 76%, with competitive results on Hi-ToM and BigToM and roughly 6x data-synthesis efficiency. The review explains ToM fundamentals, discusses applications from chatbots to negotiation AI, and notes that architecture and training matter as much as model scale.

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind — Paper Review

Paper: OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Authors: Sharmin Sultana Srishty, Kazi Mahathir Rahman, Malaika Parizat Sakkhi arXiv: 2505.10250 | cs.AI

Background: Theory of Mind

Theory of Mind (ToM) is the ability to reason about what others know, believe, and want. From the classic false-belief task (where a child must reason that someone will search a drawer because they did not see an object being moved) to high-order nested reasoning ("I know that you know that I know..."), ToM underpins human social intelligence. Large language models handle simple ToM tasks via pattern matching but collapse on genuinely novel, deeply nested social scenarios.

A key blind spot is Observer-Self Conflict: an observer who knows a fact (e.g., where a hidden gift is) while needing to act as if they do not know it. Existing benchmarks like ExploreToM, Hi-ToM, and BigToM rarely test such conflicts. On the information-asymmetry benchmark FANToM, the previous best method (ExploreToM) scored only 0.2%.

The OSCToM Approach

1. Adversarial puzzle generation: An LLM-based generator deliberately creates ToM problems the target model gets wrong — multiple agents with different information sets, recursive beliefs, and observer-self conflicts — forming an arms race where difficulty adapts as the model improves. 2. Reinforcement learning training: Rather than supervised memorization, an extended domain-specific language (DSL) structures ToM scenarios, and RL reward shaping encourages correct layered reasoning. 3. Compositional surrogate models: An 8B-parameter model outperforms far larger models by decomposing ToM into sub-skills (extracting information sets, building belief chains, detecting conflicts, synthesizing answers), improving interpretability, modularity, and efficiency.

Results

  • FANToM: 76% accuracy vs. ExploreToM's 0.2% — roughly a 380x improvement, moving from total failure to near-usable performance (humans score 90%+).
  • Hi-ToM and BigToM: remains competitive, without sacrificing generality.
  • Data efficiency: ~6x more effective data synthesis than traditional methods.

Why It Matters

ToM underlies real applications: customer-service bots detecting hidden frustration, educational AI judging whether "I get it" is genuine, medical intake uncovering unspoken details, autonomous driving predicting pedestrian intent, and negotiation agents. OSCToM shows ToM capability is not purely a function of scale — architecture and training strategy matter, good news for resource-constrained developers.

Future directions include multimodal ToM (using facial expressions and tone), dynamic ToM (updating belief models during live interaction), and ethics: an AI that models beliefs also gains the capacity to manipulate them.

References

1. Srishty, Rahman & Sakkhi (2025). OSCToM. arXiv:2505.10250. 2. Premack & Woodruff (1978). Does the chimpanzee have a theory of mind? 3. Baron-Cohen (1995). Mindblindness. MIT Press. 4. Wilf et al. (2024). FANToM. ACL. 5. Gandhi et al. (2024). ExploreToM. 6. Vaccaro & Curhan (2025). Personality Engineering with AI Agents.

Tags

#theory-of-mind#llm#reinforcement-learning#adversarial-data-generation#fanTom-benchmark#ai-reasoning#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620566