From Task Distributions to Expected Paths Lengths Distributions: Value Function Initialization in Sparse Reward Environments for Lifelong Reinforcement Learning

Soumia Mehimeh
Xianglong Tang

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

This paper studies value function transfer within reinforcement learning frameworks, focusing on tasks continuously assigned to an agent through a probabilistic distribution. Specifically, we focus on environments characterized by sparse rewards with a terminal goal. Initially, we propose and theoretically demonstrate that the distribution of the computed value function from such environments, whether in cases where the goals or the dynamics are changing across tasks, can be reformulated as the distribution of the number of steps to the goal generated by their optimal policies, which we name expected optimal path length. To test our propositions, we hypothesize that the distribution of the expected optimal path lengths resulting from the task distribution is normal. This claim leads us to propose that if the distribution is normal, then the distribution of the value function follows a log-normal pattern. Leveraging this insight, we introduce "LogQInit" as a novel value function transfer method, based on the properties of log-normality. Finally, we run experiments on a scenario of goals and dynamics distributions, validate our proposition by providing an a dequate analysis of the results, and demonstrate that LogQInit outperforms existing methods of value function initialization, policy transfer, and reward shaping.

Version published to 10.20944/preprints202502.0392.v1
Feb 6, 2025

Mastering Reinforcement Learning: Foundations, Algorithms, and Real-World Applications

This article has 16 authors:
1. Xinyuan Song
2. Keyu Chen
3. Ziqian Bi
4. Qian Niu
5. Junyu Liu
6. Benji Peng
7. Sen Zhang
8. Ming Liu
9. Ming Li
10. Xuanhe Pan
11. Jiawei Xu
12. Jinlang Wang
13. Pohsun Feng
14. Zichen Yuan
15. Li Zhang
16. Yan Zhong
This article has no evaluationsLatest version Feb 20, 2025
Introduction to Reinforcement Learning from Human Feedback: A Review of Current Developments

This article has 1 author:
1. Satyadhar Joshi
This article has no evaluationsLatest version Mar 17, 2025
Probabilistic forecasting guides dynamic decisions

This article has 3 authors:
1. Shuze Liu
2. Yang Xiang
3. Samuel J. Gershman
This article has no evaluationsLatest version Mar 21, 2025

Listed in

Abstract

Article activity feed

Related articles

Mastering Reinforcement Learning: Foundations, Algorithms, and Real-World Applications

Introduction to Reinforcement Learning from Human Feedback: A Review of Current Developments

Probabilistic forecasting guides dynamic decisions