Merge-Adversarial Training Improves Watermark Durability in Open-Source LLMs
A new arXiv preprint introduces Merge-Adversarial Training, a method designed to make watermarks in open-source large language models (LLMs) more robust to removal by model merging—a common post-training modification. The approach reportedly increases watermark detection rates by up to 51 percentage points at a 1% false positive rate, without degrading model performance. The study also evaluates watermark durability across several realistic model merging scenarios.
Why it matters: This work addresses a key challenge in tracing the origins of text generated by open-source LLMs, even after models are merged or modified post-release.
Full story at: arXiv Computation and Language ↗