Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning

Dong Liu
Yanxuan Yu
Ying Nian Wu

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

The success of large language models (LLMs) across diverse NLP tasks has elevated the importance of reasoning chain optimization as a critical step in aligning model behavior with task objectives. Existing reasoning chain tuning methods often rely on black-box heuristics or gradient-free search, which lack interpretability, generalization, and sample efficiency. In this work, we introduce Thoughts-as-Planning , a novel framework that formalizes reasoning chain optimization as a sequential decision-making process over a latent semantic space. We model the LLM as a partially observable environment and learn a latent world model that simulates the effect of reasoning chain edits on downstream outputs. A proximity-preserving embedding space is constructed to encode reasoning chain-response dynamics, enabling planning via gradient descent or reinforcement learning. Our method supports multi-scale abstraction, allowing reasoning chain edits at token, segment, and instruction levels to be integrated into a unified planner. Through extensive experiments on language understanding and generation tasks, we demonstrate that Thoughts-as-Planning outperforms state-of-the-art reasoning chain tuning baselines in efficiency, robustness, and generalization, while offering interpretability through its structured planning trajectory. Our code is available at https://github.com/FastLM/Thoughts-as-Planning .

Version published to 10.64898/2026.05.10.724161 on bioRxiv
May 15, 2026

Build on Priors: Vision-Language-Guided Neuro-Symbolic Imitation Learning for Data-Efficient Real-World Robot Manipulation

This article has 6 authors:
1. Pierrick Lorang
2. Johannes Huemer
3. Timothy Duggan
4. Kai Goebel
5. Patrik Zips
6. Matthias Scheutz
This article has no evaluationsLatest version Apr 7, 2026
A Hierarchical Multi-Agent Reinforcement Learning Framework with High-Level Guidance from Large Language Models

This article has 10 authors:
1. Jinyin Bai
2. Wei Zhu
3. Xiangchen Wang
4. KaiYang Kou
5. Shiluo Guo
6. Shuhong Liu
7. Dong Li
8. Tianjin Ni
9. Jinji Zhou
10. Yihao Zhong
This article has no evaluationsLatest version Apr 8, 2026
Bridging LLM Reasoning and Chemical Knowledge via an Evolutionary Multi-Agent Framework for Molecular Synthesis

This article has 5 authors:
1. Yicong Chen
2. Jiahua Rao
3. Jiancong Xie
4. Youhan Sun
5. Yuedong Yang
This article has no evaluationsLatest version May 6, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Build on Priors: Vision-Language-Guided Neuro-Symbolic Imitation Learning for Data-Efficient Real-World Robot Manipulation

A Hierarchical Multi-Agent Reinforcement Learning Framework with High-Level Guidance from Large Language Models

Bridging LLM Reasoning and Chemical Knowledge via an Evolutionary Multi-Agent Framework for Molecular Synthesis