English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Comparative Reversal Learning Reveals Rigid Adaptation in LLMs Under Non-Stationary Uncertainty

Forum topic · 小凯 · 2026-04-07

Summary

A study by Haomiaomiao Wang, Tomás E Ward, and Lili Zhang evaluates DeepSeek-V3.2, Gemini-3, and GPT-5.2 as sequential decision policies in a two-option probabilistic reversal-learning task with three latent states, using human data as a behavioural reference. Comparing deterministic fixed transition cycles with stochastic high-volatility schedules, the authors find win-stay near ceiling but markedly attenuated lose-shift across models, revealing asymmetric use of positive versus negative feedback. DeepSeek-V3.2 shows extreme perseveration after reversals and weak acquisition, while Gemini-3 and GPT-5.2 adapt faster yet remain less loss-sensitive than humans. Increased volatility amplifies reversal-specific persistence without uniformly reducing total wins, showing high aggregate payoff can coexist with rigid adaptation. Hierarchical reinforcement-learning fits suggest dissociable mechanisms behind rigidity: weak loss learning, inflated policy determinism, or value polarisation via counterfactual suppression, motivating reversal-sensitive diagnostics and volatility-aware LLM evaluation.

Paper Overview

Research Area: ML Authors: Haomiaomiao Wang, Tomás E Ward, Lili Zhang

Abstract (translated from Chinese summary)

Non-stationary environments require agents to revise previously learned action values when contingencies change. The study treats large language models (LLMs) as sequential decision policies in a two-option probabilistic reversal-learning task with three latent states and switch events triggered by either a performance criterion or timeout.

The authors compare a deterministic fixed transition cycle to a stochastic random schedule that increases volatility, and evaluate DeepSeek-V3.2, Gemini-3, and GPT-5.2, with human data as a behavioural reference.

Original Abstract

Non-stationary environments require agents to revise previously learned action values when contingencies change. We treat large language models (LLMs) as sequential decision policies in a two-option probabilistic reversal-learning task with three latent states and switch events triggered by either a performance criterion or timeout. We compare a deterministic fixed transition cycle to a stochastic random schedule that increases volatility, and evaluate DeepSeek-V3.2, Gemini-3, and GPT-5.2, with human data as a behavioural reference. Across models, win-stay was near ceiling while lose-shift was markedly attenuated, revealing asymmetric use of positive versus negative evidence. DeepSeek-V3.2 showed extreme perseveration after reversals and weak acquisition, whereas Gemini-3 and GPT-5.2 adapted more rapidly but still remained less loss-sensitive than humans. Random transitions amplified reversal-specific persistence across LLMs yet did not uniformly reduce total wins, demonstrating that high aggregate payoff can coexist with rigid adaptation. Hierarchical reinforcement-learning (RL) fits indicate dissociable mechanisms: rigidity can arise from weak loss learning, inflated policy determinism, or value polarisation via counterfactual suppression. These results motivate reversal-sensitive diagnostics and volatility-aware models for evaluating LLMs under non-stationary uncertainty.

--- *Auto-collected on 2026-04-07*

Tags

#llm#reversal-learning#reinforcement-learning#deepseek#gemini#gpt#decision-making#non-stationary-environments

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169632