Language model alignment is often a trade-off: the aligned model should not drift too far from the base model while also earning higher reward. A new arXiv paper (2610.01828) formalizes this as a perturbation problem, studying the asymptotics of this trade-off when the model has memory.
The abstract lays out two conditions: the aligned distribution q must remain close in probability to the original Q, and q must achieve higher expected reward than Q. It then begins to describe 'two common' approaches, but the text cuts off there. Consequently, the specific results, methods, and conclusions are not visible from the abstract alone.
That said, the framing points to a theoretical analysis of how alignment methods behave in the limit. Readers interested in the full derivations and findings will need to consult the complete paper.