A new preprint introduces Diffusion Layer Integrated Gradients (DLIG), a token attribution method built for diffusion language models. DLIG extends the established integrated gradients technique to arbitrary layers, rather than restricting analysis to the input or output, allowing attribution to be measured throughout the model.
The method is described as temporally resolved, meaning it can track how token contributions change as generation unfolds. This is intended to make the generation dynamics of diffusion language models more visible, offering a finer-grained view of when and where particular tokens matter.
Because DLIG works across layers and over time, it could serve as a diagnostic tool for understanding and improving diffusion-based text generation. The preprint focuses on the method itself, leaving room for future work to apply it to specific model behaviors or benchmarks.