Papers
arxiv:2507.14295

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Published on Jul 18
· Submitted by Licheng Liu on Jul 22
Authors:
,
,
,
,
,
,

Abstract

Training large reasoning models with multi-turn reinforcement learning using unary feedback improves both single-turn performance and multi-turn reasoning accuracy.

AI-generated summary

Multi-turn problem solving is critical yet challenging for Large Reasoning Models (LRMs) to reflect on their reasoning and revise from feedback. Existing Reinforcement Learning (RL) methods train large reasoning models on a single-turn paradigm with verifiable rewards. However, we observe that models trained with existing RL paradigms often lose their ability to solve problems across multiple turns and struggle to revise answers based on contextual feedback, leading to repetitive responses. We ask: can LRMs learn to reflect their answers in a multi-turn context? In this work, we find that training models with multi-turn RL using only unary feedback (e.g., "Let's try again") after wrong answers can improve both single-turn performance and multi-turn reasoning. We introduce Unary Feedback as Observation (UFO) for reinforcement learning, which uses minimal yet common unary user feedback during iterative problem solving. It can be easily applied to existing single-turn RL training setups. Experimental results show that RL training with UFO keeps single-turn performance and improves multi-turn reasoning accuracy by up to 14%, enabling language models to better react to feedback in multi-turn problem solving. To further minimize the number of turns needed for a correct answer while encouraging diverse reasoning when mistakes occur, we design reward structures that guide models to produce careful and deliberate answers in each turn. Code: https://github.com/lichengliu03/unary-feedback

Community

Paper author Paper submitter

This paper proposes Unary Feedback as Observation (UFO), a simple multi-turn reinforcement learning method that helps large reasoning models reflect on mistakes and improve through minimal feedback like “Let’s try again.” UFO improves multi-turn reasoning accuracy by up to 14% while maintaining single-turn performance, enabling more deliberate and flexible problem solving.

Have you found a way to reduce the range of retrieval for their context window and expression?

Sign up or log in to comment

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2507.14295 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2507.14295 in a Space README.md to link it from this page.

Collections including this paper 4