OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind — Paper Review
Paper: OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind Authors: Sharmin Sultana Srishty, Kazi Mahathir Rahman, Malaika Parizat Sakkhi arXiv: 2505.10250 | cs.AI
Background: Theory of Mind
Theory of Mind (ToM) is the ability to reason about what others know, believe, and want. From the classic false-belief task (where a child must reason that someone will search a drawer because they did not see an object being moved) to high-order nested reasoning ("I know that you know that I know..."), ToM underpins human social intelligence. Large language models handle simple ToM tasks via pattern matching but collapse on genuinely novel, deeply nested social scenarios.
A key blind spot is Observer-Self Conflict: an observer who knows a fact (e.g., where a hidden gift is) while needing to act as if they do not know it. Existing benchmarks like ExploreToM, Hi-ToM, and BigToM rarely test such conflicts. On the information-asymmetry benchmark FANToM, the previous best method (ExploreToM) scored only 0.2%.
The OSCToM Approach
1. Adversarial puzzle generation: An LLM-based generator deliberately creates ToM problems the target model gets wrong — multiple agents with different information sets, recursive beliefs, and observer-self conflicts — forming an arms race where difficulty adapts as the model improves. 2. Reinforcement learning training: Rather than supervised memorization, an extended domain-specific language (DSL) structures ToM scenarios, and RL reward shaping encourages correct layered reasoning. 3. Compositional surrogate models: An 8B-parameter model outperforms far larger models by decomposing ToM into sub-skills (extracting information sets, building belief chains, detecting conflicts, synthesizing answers), improving interpretability, modularity, and efficiency.
Results
- FANToM: 76% accuracy vs. ExploreToM's 0.2% — roughly a 380x improvement, moving from total failure to near-usable performance (humans score 90%+).
- Hi-ToM and BigToM: remains competitive, without sacrificing generality.
- Data efficiency: ~6x more effective data synthesis than traditional methods.
Why It Matters
ToM underlies real applications: customer-service bots detecting hidden frustration, educational AI judging whether "I get it" is genuine, medical intake uncovering unspoken details, autonomous driving predicting pedestrian intent, and negotiation agents. OSCToM shows ToM capability is not purely a function of scale — architecture and training strategy matter, good news for resource-constrained developers.
Future directions include multimodal ToM (using facial expressions and tone), dynamic ToM (updating belief models during live interaction), and ethics: an AI that models beliefs also gains the capacity to manipulate them.
References
1. Srishty, Rahman & Sakkhi (2025). OSCToM. arXiv:2505.10250. 2. Premack & Woodruff (1978). Does the chimpanzee have a theory of mind? 3. Baron-Cohen (1995). Mindblindness. MIT Press. 4. Wilf et al. (2024). FANToM. ACL. 5. Gandhi et al. (2024). ExploreToM. 6. Vaccaro & Curhan (2025). Personality Engineering with AI Agents.