When a language model is compressed for deployment, the usual practice is to report a divergence metric such as KL divergence to summarize how far the compressed model has moved from the original dense model. But a new arXiv paper points out that this kind of summary may not answer the question that actually matters for deployment: how many of the dense model's decisions have flipped?

The authors frame the issue around reliance. If a deployment depends on the dense model's outputs, then a single aggregate divergence score can obscure the practical impact of compression. What is needed, they argue, is a count of decision changes—how many outputs would differ in practice—rather than a statistical distance that may not correspond to real-world behavior.

The abstract is brief and cuts off mid-sentence after "We show…", so the full method and results are not available in the source text. Still, the motivation is clear: compression evaluation may need to shift from measuring divergence to measuring decision-level agreement.