An arXiv preprint (2610.11989) introduces MetaOPD, a method for on-policy distillation. On-policy distillation trains a student model on its own generated responses, using token-level supervision from a teacher. The authors argue that applying uniform weights to all tokens ignores differences in how much each token contributes to learning.

MetaOPD instead meta-learns token weights, aiming to make distillation more effective. The abstract criticizes existing weighting methods, but the announcement text cuts off before specifying what those methods rely on. The paper's core contribution is the meta-learned weighting scheme itself.

Because the source is a single abstract, the article does not compare MetaOPD with other approaches beyond noting that it targets the limitations of uniform weighting and prior weighting methods.