Most current in-context reinforcement learning methods pretrain Transformers with supervised behavior-prediction objectives. That lets the model infer the task from context, but it also means the learned policy inherits the limitations of the training demonstrations. The new paper, posted on arXiv, identifies this dependence as a key weakness and proposes an alternative: training with Q-targets rather than behavior cloning.
Using Q-targets, the authors argue, makes the resulting policy provably effective even when the pretraining data is weak or suboptimal. The theoretical framing suggests that the policy's performance does not hinge on the quality of the demonstrations, but on the structure of the Q-learning objective. The abstract does not include experimental details, so the claims rest on the theoretical analysis presented in the paper.
The work is positioned as a step toward making in-context RL more reliable in practice, where expert demonstrations are often scarce or noisy. By shifting the training signal from mimicking behavior to estimating action values, the method aims to separate task inference from policy quality.