RANDOMNESS REUSE IN LLM WATERMARKS: SECURITY BOUNDS AND BLACK-BOX
- LecturerDr. Yu Long Chen (COSIC research group, KU Leuven)
Host: 鐘楷閔Kai-Min Chung - Time2026-10-07 (Wed.) 10:15 ~ 12:15
- LocationAuditorium 101 at IIS new Building
Abstract
Generative watermarking is increasingly used to identify LLM-generated content.
While theoretical security notions and practical attacks have been studied, their relationship and whether attacks exploit general design properties or implementation details remain unclear. We study these questions through formal analysis and black-box spoofing attacks on SynthID-Text. We first connect two watermark security goals usually analyzed separately: undetectability and spoofing resistance, showing that a successful spoofing attack also yields a distinguisher between watermarked and clean LLM generation when one terminal verification query on a fresh candidate is allowed. We then derive a general security bound for context-dependent watermarking schemes in which randomness reuse explicitly degrades security. Our experiments test whether this reuse is exploitable in practice. On SynthID-Text, we developed black-box spoofing attacks that recover watermark bias from model outputs. An implementation-specific weakness in the public Hugging Face implementation further strengthens the attack, yielding joint detection-and-automated-quality rates of 82.6% and 88.7% on Gemma-2B and Gemma-9B prompts, respectively.
While theoretical security notions and practical attacks have been studied, their relationship and whether attacks exploit general design properties or implementation details remain unclear. We study these questions through formal analysis and black-box spoofing attacks on SynthID-Text. We first connect two watermark security goals usually analyzed separately: undetectability and spoofing resistance, showing that a successful spoofing attack also yields a distinguisher between watermarked and clean LLM generation when one terminal verification query on a fresh candidate is allowed. We then derive a general security bound for context-dependent watermarking schemes in which randomness reuse explicitly degrades security. Our experiments test whether this reuse is exploitable in practice. On SynthID-Text, we developed black-box spoofing attacks that recover watermark bias from model outputs. An implementation-specific weakness in the public Hugging Face implementation further strengthens the attack, yielding joint detection-and-automated-quality rates of 82.6% and 88.7% on Gemma-2B and Gemma-9B prompts, respectively.