Transition from model-free to structure-informed decision making in dynamic environments
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Reinforcement learning theory formulates distinct decision-making strategies, including reactive model-free and deliberative model-based strategies. This study investigates how mice adjust their reinforcement learning strategies while learning decision-making in dynamic environments. Unlike previous studies that focused on behaviors after extensive training periods, we analyzed changes in learning strategies in the course of training of a two-step decision-making task with probabilistic state transition and fluctuating reward probabilities. Our statistical behavioral analysis showed that the stay-probability following common and rare transitions diverged with training, a signature of strategies that utilize knowledge of task structure. We fit various reinforcement learning strategies to behavioral data and found that structure-informed strategies became increasingly dominant in their behaviors during training. Whereas previous studies emphasized transition from goal-directed to habitual strategies after extensive training, which were often associated with model-based and model-free strategies, respectively, our results newly demonstrate a shift from model-free to structure-informed strategies in early training in mice.
Author summary
Reinforcement learning theory allows us to examine how we make decisions and what approaches we use to optimize rewards. Most previous research, however, has examined animal behavior only after extensive training. Here we analyzed how mice adjust their reinforcement learning strategies as they are trained in a two-step decision-making task. Initially, mice relied on reactive model-free strategies, but as training progressed, their behavior began to incorporate knowledge of task structure. While previous studies suggested transition from model-based to model-free strategies with extensive training, our study revealed the opposite in the early stage of training.