Paper Overview
Field: Machine Learning Authors: Noémi Éltető, Nathaniel D. Daw, Kimberly L. Stachenfeld, Kevin J. Miller Published: 2026-06-10 arXiv: 2606.12386
Abstract
Advancing scientific understanding through mechanistic modeling requires posing the right experimental questions to yield maximally informative data. To automate this pursuit within cognitive science, the authors introduce ATLAS (Active Theory Learning for Automated Science), an active learning framework for the data-driven discovery of interpretable behavioral models.
ATLAS iterates between two stages:
1. Generating mechanistic hypotheses, instantiated as a diverse ensemble of sparse neural networks (Disentangled RNNs) 2. Designing experiments that optimally distinguish between competing hypotheses
Evaluation
The framework is tested on the problem of recovering reinforcement learning agents from their behavior in bandit tasks. ATLAS designs varied sequences of qualitatively novel experiments with temporal structure tailored to the underlying agent features.
Models trained on these experiments are evaluated against a comprehensive set of metrics capturing:
- Behavioral similarity
- Structural similarity
- Computational similarity
- ATLAS achieves a 5–10x improvement in sample efficiency compared to random experiment selection across all metrics.
- Performance is further validated through comparison with expert-designed experiments from the literature.
Results
Implications
These computational results demonstrate ATLAS's potential to accelerate human-interpretable insight in cognitive science and other scientific domains that depend on discovering mechanistic models.
--- *Source: arXiv:2606.12386*