The unique value of zero prediction errors in reinforcement learning
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Updating beliefs when necessary is at the cornerstone of learning. A fundamental problem is to describe under which conditions humans update their internal models of the world. The general assumption across animal, human, and artificial learning models 1–3 has been that updates occur when outcomes deviate from expectations. Perfectly predicted outcomes cause no learning. As a result, no research has examined cases in which predictions are surprisingly perfect. Here, we empirically test this assumption and find that zero prediction errors are psychologically unique. We show that after zero prediction errors, momentary happiness is the highest, and belief updates do indeed occur in a pattern that cannot be reproduced by a benchmark model. We present a new model that captures this non-linear pattern in belief updating by postulating that zero prediction errors elicit a distinct latent belief state, guiding subsequent updating. This latent state then tracks neural activity patterns measured with EEG precisely when zero prediction errors occur, exactly as the model would predict. Crucially, the strength of the neural activity during this time window exhibits a dissociation in predicting the next belief update depending on whether feedback was a zero prediction error or a regular prediction error. Overall, we provide strong evidence that surprisingly perfect predictions are treated in a unique, non-linear fashion at affective, behavioral and neural levels. Being surprisingly accurate can function as a distinct belief updating signal, conforming trial-and-hit learning.