A new preprint on arXiv, CM-DPO, takes aim at a blind spot in Direct Preference Optimization (DPO) for LLM planning. The authors note that standard DPO treats every constraint violation as equally bad, so a small budget overshoot and a massive one produce the same training signal. That loses important information about how far a plan strays from the stated requirements.
The proposed approach, Constraint-Margin DPO, adds a margin that reflects the size of the violation, allowing the optimization to distinguish between minor and severe failures. The abstract also warns that DPO is susceptible to length and style bias when preference pairs are constructed, and CM-DPO appears designed to mitigate that as well.
Because the abstract is brief, the preprint does not yet provide experimental details in the available text. Still, the core idea is clear: for planning tasks, preference learning should care not just whether a constraint was broken, but by how much.