Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
TL;DR AI
2 min readKey summary
Researchers found that averaging probability outputs from multiple LLMs can erase watermark signals, causing detection to fail.
They introduce WASH to reconcile differences in vocabularies and tokenization across models, enabling the attack in realistic settings.
Across six watermarking methods and three LLMs, 3-model ensembles pushed detection scores below common thresholds and reduced true positive rates.
The study shows a major weakness in AI text watermarking, with ensemble averaging sometimes improving output quality and efficiency at the same time.
