Inferring learning rules during de novo task learning

Victor Geadah
Jonathan W. Pillow

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Identifying the learning rules that govern behavior is a central problem in neuroscience. While reinforcement learning (RL) offers a unifying theoretical framework, most empirical studies of animal learning behavior have focused on non-stationary environments (e.g. changing reward probabilities in a known task), as opposed to acquiring an entirely new task from scratch. Here we introduce a statistical framework to infer reinforcement learning rules directly from single-animal behavior. Applied to mice learning a perceptual decision-making task, our approach reveals that policy-gradient–like rules capture de novo task learning better than classical temporal-difference algorithms. By fitting flexible parametric learning rules, we uncover systematic deviations from standard RL models, including side-specific learning rates and negative reward baselines. Together, these parameters account for side-biased learning, as well as forgetting and consecutive errors due to aversive responses to incorrect trials. Extending the framework with latent, dynamic learning rates further reveals that animals adapt their learning rates over training and across curricula. These results provide a statistical account of how animals learn from scratch and highlight key departures from classical reinforcement learning algorithms.

Version published to 10.1101/2025.09.29.679295 on bioRxiv
Sep 30, 2025

A Brief Tutorial on Reinforcement Learning: From MDP to DDPG

This article has 1 author:
1. Tian Zhang
This article has no evaluationsLatest version Jan 6, 2026
Learning from oneself: Function learning with self-generated samples

This article has 2 authors:
1. Hidehito Honda
2. Rina Kagawa
This article has no evaluationsLatest version Jan 14, 2026
Delayed reward information is underweighted in reinforcement learning with dispersed feedback

This article has 3 authors:
1. Miruna Cotet
2. David Poensgen
3. Ian Krajbich
This article has no evaluationsLatest version Jan 9, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

A Brief Tutorial on Reinforcement Learning: From MDP to DDPG

Learning from oneself: Function learning with self-generated samples

Delayed reward information is underweighted in reinforcement learning with dispersed feedback