Liquid AI has released LFM2.5-DSpark, a new version of its large foundation model that is designed to run inference much faster. According to a Hugging Face blog post, the model delivers up to 3.2x faster inference compared to the original LFM2.5. The improvement is attributed to a sparse attention mechanism that skips unnecessary computations, particularly for long-context tasks.

The blog does not include detailed benchmark methodology or independent test results, so the exact conditions for the 3.2x figure are unclear. However, the claim suggests that the model is optimized for real-time or high-throughput applications where latency matters. The post also notes that the model retains the same architecture and capabilities as LFM2.5, meaning the speedup does not come at the cost of accuracy.

This release is part of a broader trend in AI research toward more efficient inference, as models grow larger and deployment costs rise. Sparse attention is one of several techniques being explored to reduce the computational burden of processing long documents or conversations. Whether the 3.2x speedup holds in practice will depend on hardware, sequence length, and other factors not detailed in the source.