Recent advances in Image-to-Video generation allow a single image to be animated into a convincing
video under text guidance, raising serious copyright and privacy risks. We propose
Anti-Prompt, an image protection approach that injects imperceptible perturbations
into an image, inducing visible inconsistencies and structural failures in text-guided I2V generation.
Our method is motivated by a simple empirical observation: when text guidance is removed from modern
I2V models, generation quality degrades markedly, not only in motion realism but also in subject
preservation, structural coherence, and temporal consistency.
Building on this insight, Anti-Prompt attenuates text-conditioned interactions during denoising while
strengthening visual-only pathways. To evaluate protection behavior, we also introduce a Video-LLM
protocol that scores subject preservation, structural consistency, dynamic consistency, and artifact
suppression using frame-grounded observations.
01Architecture-aware protection for full- and cross-attention I2V models
02Efficient optimization without auxiliary reference-video generation
03Frame-grounded evaluation focused on protection-induced failures