A new preprint on arXiv introduces ReHoPER, a method designed to improve reasoning in large language models without any fine-tuning. The approach is inference-only and zero-shot, meaning it works on pre-trained models at generation time. Rather than asking the model to answer a question directly, ReHoPER plans a sequence of intermediate questions and answers along multiple paths before committing to a final response.

The core idea comes from receding-horizon planning, a technique common in control theory. Instead of planning the entire reasoning chain in one step, the model repeatedly looks ahead a limited number of steps, answers intermediate questions, and then updates its plan. This iterative process lets the model explore several possible reasoning paths and use the results to guide its final answer.

The paper is a single source and has not yet been peer-reviewed, so its claims should be treated as preliminary. Still, the method's simplicity is notable: it requires no extra training data, no fine-tuning, and no architectural changes. If the results hold up, ReHoPER could offer a lightweight way to boost reasoning performance on existing models.