Modelling dopaminergic signals associated with habit formation through temporal-difference action learning
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Action-selection is determined by a combination of goal-directed and habitual processes. Habits are defined as the reward-independent, stimulus-response relationships which form when an action is regularly executed in the same context, regardless of outcome. An influential computational model proposes that habit formation is driven by action prediction errors which occur when non-habitual actions are taken. It has been further suggested that action prediction errors are encoded in activity of specific dopamine neurons, and it has been recently observed that dopamine activity in the tail of the striatum follows a pattern consistent with the action prediction errors. However, the original models capture changes in habits across trials, but do not describe the time-course of action prediction errors within trials, hence it is difficult to directly compare them with dopamine activity. We begin by outlining the ‘temporal-difference action learning’ algorithm, which uses biologically-plausible mechanisms to determine how dynamic changes in action intensity influence the resultant prediction errors across near-continuous time. We then demonstrate that dopaminergic data recently collected from the tail of the striatum is better represented by action prediction errors than reward prediction errors. Overall, our results support the existence of value-free action prediction errors and associated habitual behaviour in dopaminergic signals.
Author summary
Whenever we choose one action over another, there are two ways that the selection can be made. We could take the time to consider what we want to achieve, calculate which action is the most likely to give us that outcome and balance it against the possible negative consequences. These ‘goal-directed’ calculations are very time-consuming and our brains could not possibly do it for every choice. Instead, we often rely on the second method, ‘habits’, which learn to copy the actions that were most often chosen in the past. In this paper, we present a new model of learning that is based on biologically plausible brain networks and applies action prediction errors to update our ‘habits’ across continuous time. Using simulations, we reveal testable predictions that are specific to our ‘temporal-difference action learning’ model and build an intuition for its behaviour. Finally, this model is tested against real dopaminergic data from the tail of the striatum, and we show that our model provides better explanation for these data, than classic ‘reward-based’ reinforcement learning models.