A new open-weight model called Saluki 27B shows that extreme quantization does not always mean a uniform drop in capability. The model is a 2-bit GGUF version of Qwen3.8-27B, compressed to just 7.89 GB and released under the Apache 2.0 license. That makes it roughly one-seventh the size of the original 54 GB checkpoint.
Despite the aggressive compression, Saluki 27B beats the original Qwen3.8-27B on tool calling, a task that requires the model to correctly select and invoke external functions. The trade-off is that it gives up ground on competition math and reasoning, where the full-size model remains stronger. This suggests that quantization can preserve or even enhance certain skills while degrading others.
The result is notable for practical deployment: a small, locally runnable model can handle tool-calling workloads competitively, even if it is not the best choice for math-heavy reasoning. The release also highlights how benchmark-specific quantization outcomes can be, and the importance of testing compressed models on the actual tasks they will be used for.