Reinforcement Learning with Verifiable Rewards (RLVR) has become a common way to improve the reasoning abilities of language models. However, as the abstract of a new arXiv paper notes, RLVR does not explicitly account for calibration during training. This means that even if a model reasons correctly, its confidence may not reflect the true likelihood of being right.

The paper, titled RL-ARC, proposes a calibration approach guided by the model's own reasoning. The idea is to use uncertainty signals derived from the reasoning process to adjust predictions, rather than relying solely on the final answer. This is intended to produce models that are both strong at reasoning and well calibrated.

The abstract is brief, so the exact mechanism and experimental results are not detailed here. The paper is listed as a new announcement on arXiv, and the full text would be needed to evaluate the method's effectiveness.