A new arXiv preprint tackles a safety gap in text-to-video (T2V) diffusion models: their ability to generate realistic depictions of violent actions like kicking, stabbing, and shooting. The authors argue that these capabilities raise concerns that motivate targeted concept erasure, a technique previously explored mostly for static images.
The work focuses on "motion concept unlearning," aiming to remove specific actions from a model's generative repertoire. This is distinct from erasing objects or visual styles, as it targets the temporal dynamics of a scene.
However, the abstract provided is truncated mid-sentence, ending at "Although concept erasure has been…". Consequently, the specific unlearning approach, the models tested, and the effectiveness of the method are not detailed in the available text. Readers should consult the full paper for those results.