Knowledge distillation is a common approach to compressing large language models: a smaller student model is trained to absorb the factual knowledge of a larger teacher, making deployment more efficient.

The new arXiv paper highlights a directional failure mode in this process. A teacher may recall a relation in one direction while failing to generate the answer when the same relation is probed in the opposite direction. This asymmetry means a student trained on one-directional outputs can inherit gaps that a straightforward accuracy check would miss.

Titled "Distilling Directional Verification," the paper addresses this problem directly. The abstract provided is truncated, so the proposed method and experimental results are not described in the source; the title indicates that verification of directional knowledge is the intended remedy.