Impaired Reinforcement Learning Underlying Explore-Exploit Decision Making in Theft Recidivists

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Larceny imposes profound societal and economic burdens; however, punitive judicial measures frequently fail to deter recidivism. The neurobehavioral mechanisms driving habitual offending, whether instrumental or kleptomanic, in theft recidivists remain poorly understood. In this study, we investigated explore-exploit decision-making and underlying reinforcement learning architectures in theft recidivists with a 4-arm bandit task while prefrontal cortex (PFC) hemodynamics were continuously monitored using functional near-infrared spectroscopy (fNIRS). Model-free behavioral analyses revealed that non-kleptomanic (TR-K) but not kleptomanic (TR+K) theft recidivists accumulated significantly higher cumulative regret and made fewer optimal choices compared to control individuals without criminal records (CT). Model-based analyses through Variational Bayesian Analysis identified Q-learning with decay model as the optimal computational fit. Parameter extraction using the model demonstrated that the TR-K group exhibited a lower learning rate than CT and TR+K groups, indicating a learning deficit in updating action values following environmental feedback. fNIRS tracking of trial-by-trial latent reinforcement variables revealed that while PFC activity was modulated by these variables, group differences were characterized by static baseline hemodynamic shifts rather than rewirings of value-tracking neural circuits. These findings suggest that non-kleptomanic recurrent thefts may be associated with impaired reinforcement learning mechanisms, challenging current punitive deterrence models.

SIGNIFICANCE STATEMENT

This study shows a neurobehavioral mechanism underlying recurrent theft, challenging the traditional assumption that recidivism stems from rational choice or moral failure. By utilizing a reinforcement learning framework, this study demonstrates that instrumental, but not kleptomanic, theft recidivists possess a diminished learning rate. Such learning deficit prevents them from efficiently updating the expected value of their actions following feedback, such as incarceration. Consequently, standard punitive deterrence models fail for this specific population. These findings highlight the need to pivot justice systems toward neurobehavioral rehabilitation strategies that target reinforcement learning deficits.

Article activity feed