Overview
- Field: Neuroscience / Reinforcement Learning / Dopamine
- Authors: Sheng Gong, Alyssa Martell, Joshua T. Dudman, Luke T. Coddington
- Institution: HHMI Janelia Research Campus
- Published: 2026-05-21, Science (Volume 392, Issue 6800)
- DOI: 10.1126/science.aeb0813
- Overturning a traditional assumption: Decades of neuroscience assumed learning speed depends mainly on the number of experiences (repetitions), independent of reward size.
- The "cookie vs. M&M" effect: Thirsty mice mastered a task within one day using a few large rewards (100μL), whereas thousands of small rewards (5μL) took weeks.
- Eliminating individual differences: With large rewards, all subjects reached expert level within days; with small-reward protocols, learning times varied widely (from one week to one month).
- Dopamine mechanism: Large rewards not only produce larger dopamine peaks—more critically, they extend the duration of the dopamine signal.
- Three learning components: Increased per-repetition learning efficiency, enhanced cross-day memory retention, and increased task engagement—with engagement being the largest determinant.
- Optogenetic validation: Artificially prolonging the dopamine signal associated with small rewards successfully reproduced most of the learning benefits of large rewards.
Abstract
Standard animal learning studies minimize individual rewards to maximize the number of reinforcement episodes. This work examined how reward size affects initial learning in naive mice across five behavioral paradigms. Notably, large rewards significantly improved learning efficiency through dissociable effects on within-session learning, cross-session learning, and task engagement. The duration and amplitude of ventral striatal dopamine release scaled proportionally with reward size, and prolonged optogenetic enhancement of dopamine reward responses reproduced most—but not all—of the learning benefits of very large rewards. These findings suggest the reinforcement learning efficiency of animals has traditionally been underestimated, and dopamine signaling of reward scales task engagement with absolute reward magnitude.
Key Findings
Implications for AI
The paper has deep implications for reinforcement learning in AI: sparse or small-reward schemes commonly used in RL experiments may severely underestimate a system's learning potential. Larger reward signals (longer dopamine waves) imply higher learning rates and faster convergence.
--- *Auto-collected on 2026-07-03*