Mice adapt their learning rate to stochasticity and volatility through a simple heuristic

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Learning from outcomes requires weighing each surprise by how much it is likely to mean. Two properties of the environment set that weight in opposite directions: volatility, the rate at which contingencies change, and stochasticity, the noise in outcomes around a stable contingency. Whether animals separate the two is unresolved, in part because varying reward probability in a reversal task makes outcomes noisier and the two options harder to tell apart at once. We trained male and female mice on a probabilistic reversal task for intracranial self-stimulation, scaling reward magnitude with reward probability so that the expected reward of each option and the difference between them stayed constant while outcome variance varied, and crossing this with different reversal rates. Simulations show that adapting the learning rate alone and adapting the decision policy alone are near-equivalent solutions across the environments we tested. Mice adopted the first adjustment: their fitted learning rate rose with volatility and fell with stochasticity, in both directions predicted by reward maximization, while their inverse temperature stayed fixed. This joint dependence was reproduced only by a meta-learning rule that reads stochasticity from the size of prediction errors and volatility from how often recent outcomes favored switching.

Article activity feed