Overview
Digital watermarking serves as a mechanism for source attribution in AI-generated images. The efficacy of this mechanism is contingent on its resilience to removal attacks. Some existing removal attacks operate by attempting to modify the image such that the decoded watermark diverges from its original form. However, this approach can inadvertently lead to the generation of an 'inverted watermark,' which remains detectable, thus failing the removal objective. Subsequent efforts to alter such a watermark risk further degradation of image quality without achieving successful removal.
To address these limitations, a novel technique named LiBRA (Latent In-band Bidirectional Removal Attack) has been developed. LiBRA's primary objective is to render watermarks undetectable while simultaneously maintaining the visual quality of the image. Its methodology deviates from previous attacks by adjusting the image to conceal the watermark without inducing additional changes that might compromise image quality. This contrasts with approaches that continuously push decoded bits away from the original watermark, even when such changes do not prevent detectability and instead damage the image.
Research Context
The reliability of digital watermarking for attributing AI-generated images is directly linked to its resistance against removal attacks. Conventional removal strategies often focus on forcing a difference between the decoded watermark and the original. A significant challenge with these strategies is the potential for creating an inverted watermark. An inverted watermark, while altered, remains detectable, signifying a failed removal attempt. Continuous manipulation to correct such an inversion often results in unnecessary image degradation, highlighting a gap in effective and quality-preserving watermark removal techniques.
Approach
LiBRA operates with access to the watermark key and the decoder. It performs bounded modifications within the latent space of a public autoencoder. A core distinction of LiBRA from inversion-driven objectives is its guidance mechanism for average decoding confidence. Instead of perpetually pushing the watermark towards inversion, LiBRA directs this confidence toward a state of random guessing. This bidirectional optimization helps in preventing the outcome of an inverted yet still detectable watermark, a common failure mode for other attacks.
The flexibility afforded to individual bits during this process allows for image-quality constraints to prioritize changes that inflict less damage. Furthermore, an optional frequency-guided mask can be employed to restrict the location of these changes. The verification of watermark removal by LiBRA is conducted using an exact two-sided binomial test. This statistical method is utilized to confirm successful removal, rather than relying on the assumption that merely reaching the confidence target guarantees success.
Findings
- LiBRA adjusts the image to conceal watermarks without encouraging further changes that could degrade image quality, contrasting with attacks that continuously push decoded bits away from the original watermark even when detectability is preserved.
- With access to the watermark key and decoder, LiBRA makes bounded changes in a public autoencoder's latent space.
- Unlike inversion-driven objectives, LiBRA guides average decoding confidence toward random guessing from either direction, which helps avoid an inverted but detectable watermark.
- Leaving individual bits flexible allows image-quality constraints to favor less damaging changes.
- An optional frequency-guided mask limits the location of changes made by LiBRA.
- Removal verification is performed using an exact two-sided binomial test, rather than assuming the confidence target guarantees success.