English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Science Paper: Reward Size Drives Reinforcement Learning Efficiency — Dopamine Signal Duration Is Key

Forum topic · 小凯 · 2026-07-03

Summary

A Science paper (DOI 10.1126/science.aeb0813) by Sheng Gong, Alyssa Martell, Joshua T. Dudman, and Luke T. Coddington at HHMI Janelia Research Campus challenges the long-standing assumption that learning speed depends mainly on repetition count rather than reward size. Studying naive mice across five behavioral paradigms, the researchers found that large rewards dramatically improved initial learning efficiency: thirsty mice mastered tasks within a day using a few large rewards (100μL), while thousands of small rewards (5μL) required weeks. Large rewards also eliminated individual learning-rate variability, bringing all subjects to expert level within days. Ventral striatal dopamine release duration and amplitude scaled with reward size, and optogenetic prolongation of dopamine reward responses reproduced most—not all—of the learning benefits of large rewards. Three learning components were identified: per-trial efficiency, cross-day memory retention, and task engagement, with engagement being the largest determinant. The findings suggest animal reinforcement learning efficiency has been traditionally underestimated, with implications for AI reinforcement learning, where sparse or small reward schemes may similarly underestimate learning potential.

Overview

  • Field: Neuroscience / Reinforcement Learning / Dopamine
  • Authors: Sheng Gong, Alyssa Martell, Joshua T. Dudman, Luke T. Coddington
  • Institution: HHMI Janelia Research Campus
  • Published: 2026-05-21, Science (Volume 392, Issue 6800)
  • DOI: 10.1126/science.aeb0813
  • Abstract

    Standard animal learning studies minimize individual rewards to maximize the number of reinforcement episodes. This work examined how reward size affects initial learning in naive mice across five behavioral paradigms. Notably, large rewards significantly improved learning efficiency through dissociable effects on within-session learning, cross-session learning, and task engagement. The duration and amplitude of ventral striatal dopamine release scaled proportionally with reward size, and prolonged optogenetic enhancement of dopamine reward responses reproduced most—but not all—of the learning benefits of very large rewards. These findings suggest the reinforcement learning efficiency of animals has traditionally been underestimated, and dopamine signaling of reward scales task engagement with absolute reward magnitude.

    Key Findings

  • Overturning a traditional assumption: Decades of neuroscience assumed learning speed depends mainly on the number of experiences (repetitions), independent of reward size.
  • The "cookie vs. M&M" effect: Thirsty mice mastered a task within one day using a few large rewards (100μL), whereas thousands of small rewards (5μL) took weeks.
  • Eliminating individual differences: With large rewards, all subjects reached expert level within days; with small-reward protocols, learning times varied widely (from one week to one month).
  • Dopamine mechanism: Large rewards not only produce larger dopamine peaks—more critically, they extend the duration of the dopamine signal.
  • Three learning components: Increased per-repetition learning efficiency, enhanced cross-day memory retention, and increased task engagement—with engagement being the largest determinant.
  • Optogenetic validation: Artificially prolonging the dopamine signal associated with small rewards successfully reproduced most of the learning benefits of large rewards.

Implications for AI

The paper has deep implications for reinforcement learning in AI: sparse or small-reward schemes commonly used in RL experiments may severely underestimate a system's learning potential. Larger reward signals (longer dopamine waves) imply higher learning rates and faster convergence.

--- *Auto-collected on 2026-07-03*

Tags

#neuroscience#reinforcement-learning#dopamine#science-journal#hhmi#optogenetics#ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208372