Extreme low-bit compression of large language models becomes especially hard when weights, activations, and KV caches are quantized together. Each component has a different distribution, and quantization errors can compound as they propagate through the network. This makes naive joint quantization impractical.
A new paper on arXiv proposes CanonQ, a method designed specifically for this setting. While the abstract does not detail the inner workings, it frames CanonQ as a response to the interaction of errors across layers and the need to handle heterogeneous distributions simultaneously.
The work targets an aggressive W2A4KV2 configuration—2-bit weights, 4-bit activations, and 2-bit KV caches—pushing beyond typical mixed-precision schemes. The authors argue that addressing the joint problem is key to making such extreme compression viable, though the full technical approach remains to be seen in the complete paper.