Mice exhibit stochastic and efficient action switching during probabilistic decision making

Celia C Beron; Shay Q Neufeld; Scott W Linderman; Bernardo L Sabatini

doi:10.1073/pnas.2113961119

Mice exhibit stochastic and efficient action switching during probabilistic decision making

Proc Natl Acad Sci U S A. 2022 Apr 12;119(15):e2113961119. doi: 10.1073/pnas.2113961119. Epub 2022 Apr 6.

Authors

Celia C Beron^{1

2}, Shay Q Neufeld^{1

2}, Scott W Linderman^{3

4}, Bernardo L Sabatini^{1

2}

Affiliations

¹ Department of Neurobiology, Harvard Medical School, Boston, MA 02115.
² HHMI, Harvard Medical School, Boston, MA 02115.
³ Department of Statistics, Stanford University, Stanford, CA 94305.
⁴ Wu Tsai Neurosciences Institute, Stanford University, Stanford, CA 94305.

Abstract

In probabilistic and nonstationary environments, individuals must use internal and external cues to flexibly make decisions that lead to desirable outcomes. To gain insight into the process by which animals choose between actions, we trained mice in a task with time-varying reward probabilities. In our implementation of such a two-armed bandit task, thirsty mice use information about recent action and action–outcome histories to choose between two ports that deliver water probabilistically. Here we comprehensively modeled choice behavior in this task, including the trial-to-trial changes in port selection, i.e., action switching behavior. We find that mouse behavior is, at times, deterministic and, at others, apparently stochastic. The behavior deviates from that of a theoretically optimal agent performing Bayesian inference in a hidden Markov model (HMM). We formulate a set of models based on logistic regression, reinforcement learning, and sticky Bayesian inference that we demonstrate are mathematically equivalent and that accurately describe mouse behavior. The switching behavior of mice in the task is captured in each model by a stochastic action policy, a history-dependent representation of action value, and a tendency to repeat actions despite incoming evidence. The models parsimoniously capture behavior across different environmental conditionals by varying the stickiness parameter, and like the mice, they achieve nearly maximal reward rates. These results indicate that mouse behavior reaches near-maximal performance with reduced action switching and can be described by a set of equivalent models with a small number of relatively fixed parameters.

Keywords: Bayesian inference; decision making; explore–exploit; perseveration; stochastic choice.

MeSH terms

Animals
Choice Behavior*
Decision Making*
Mice* / psychology
Reward
Uncertainty

Abstract

MeSH terms

Grants and funding