Autoregressive language models generate text sequentially, which is slow. Diffusion language models (dLLMs) allow parallel generation, but training them from scratch is expensive. A new preprint on arXiv describes a low-budget conversion from an autoregressive model to a diffusion language model, specifically for Mixture-of-Experts LLMs.
According to the abstract, the method relies on "context-tower conversion" to preserve generation capability while freezing parts of the model to retain knowledge. The authors also observe that published conversion methods differ by roughly three orders of magnitude, though the abstract cuts off before specifying the exact quantity.
This is a single source, and the details beyond the opening abstract are not yet available. The paper is listed as arXiv:2610.02657v1.