Comparative Reversal Learning Reveals Rigid Adaptation in LLMs Under Non-Stationary Uncertainty
Forum topic · 小凯 · 2026-04-07
Summary
This paper evaluates how large language models adapt to non-stationary uncertainty using a two-option probabilistic reversal-learning task with three latent states and switch events triggered by performance criteria or timeouts. Testing DeepSeek-V3.2, Gemini-3, and GPT-5.2 against human behavioral data, the authors find that all models show win-stay behavior near ceiling while lose-shift behavior is markedly attenuated, indicating asymmetric use of positive versus negative evidence. DeepSeek-V3.2 exhibits extreme perseveration after reversals and weak acquisition, while Gemini-3 and GPT-5.2 adapt faster but remain less loss-sensitive than humans. Random, higher-volatility transition schedules amplified reversal-specific persistence across LLMs without uniformly reducing total wins, showing that high aggregate payoff can coexist with rigid adaptation. Hierarchical reinforcement-learning fits identify dissociable mechanisms of rigidity: weak loss learning, inflated policy determinism, or value polarisation via counterfactual suppression. The results motivate reversal-sensitive diagnostics and volatility-aware evaluation for LLMs.
Overview
- Field: Machine Learning
- Authors: Haomiaomiao Wang, Tomas E Ward, Lili Zhang
Abstract (original)
Non-stationary environments require agents to revise previously learned action values when contingencies change. We treat large language models (LLMs) as sequential decision policies in a two-option probabilistic reversal-learning task with three latent states and switch events triggered by either a performance criterion or timeout. We compare a deterministic fixed transition cycle to a stochastic random schedule that increases volatility, and evaluate DeepSeek-V3.2, Gemini-3, and GPT-5.2, with human data as a behavioural reference. Across models, win-stay was near ceiling while lose-shift was markedly attenuated, revealing asymmetric use of positive versus negative evidence. DeepSeek-V3.2 showed extreme perseveration after reversals and weak acquisition, whereas Gemini-3 and GPT-5.2 adapted more rapidly but still remained less loss-sensitive than humans. Random transitions amplified reversal-specific persistence across LLMs yet did not uniformly reduce total wins, demonstrating that high aggregate payoff can coexist with rigid adaptation. Hierarchical reinforcement-learning (RL) fits indicate dissociable mechanisms: rigidity can arise from weak loss learning, inflated policy determinism, or value polarisation via counterfactual suppression. These results motivate reversal-sensitive diagnostics and volatility-aware models for evaluating LLMs under non-stationary uncertainty.Key Points (translated summary)
- The study assesses LLM adaptability in non-stationary environments via a probabilistic reversal-learning task.
- All models show a "win-stay / lose-shift" asymmetry: positive evidence is used near ceiling, while negative evidence is used much more weakly.
- DeepSeek-V3.2 shows extreme perseveration after reversals and weak acquisition; Gemini-3 and GPT-5.2 adapt faster but are still less loss-sensitive than humans.
- High total payoff can coexist with rigid adaptation.
- Rigidity may stem from weak loss learning, over-confident policy determinism, or value polarisation caused by counterfactual suppression.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169611