Compressing a large language model can slash inference costs and memory footprint, but deciding which compression method to use is still largely trial and error. The new paper argues that this is because comparable resource reductions—similar cuts to model size or compute—can produce very different drops in capability, making it hard to predict the cost of a given compression choice.
To address this, the authors introduce capability scaling-down laws for LLM compression. These laws aim to model how a model's capabilities degrade as it is compressed, providing a systematic way to anticipate the trade-off between resource savings and capability loss. The work is positioned as a step toward making compression choices less empirical and more principled.