A new preprint on arXiv introduces LiBRA, a method for removing watermarks from AI-generated images. Unlike earlier attacks that only tweak pixel values, LiBRA works in the latent space of a generative model, optimizing the image representation to disrupt the embedded watermark. This makes the removal more effective and less detectable.

The key innovation is its bidirectional optimization. LiBRA not only forces the decoded watermark to differ from the original but also ensures that the image is no longer flagged as watermarked by a detector. This dual objective makes the attack more practical, as it can evade both watermark verification and forensic detection.

The paper reports that LiBRA outperforms existing removal methods in terms of both watermark removal success and image quality preservation. However, the authors note that their method assumes access to the watermark decoder, which may not always be available in real-world scenarios. This limitation is acknowledged in the preprint, but the method still represents a significant step in understanding the vulnerabilities of current watermarking systems.