Switch language한국어
Back to the list

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

TL;DR AI

Key summary

2 min read
  1. Researchers found that averaging probability outputs from multiple LLMs can erase watermark signals, causing detection to fail.

  2. They introduce WASH to reconcile differences in vocabularies and tokenization across models, enabling the attack in realistic settings.

  3. Across six watermarking methods and three LLMs, 3-model ensembles pushed detection scores below common thresholds and reduced true positive rates.

  4. The study shows a major weakness in AI text watermarking, with ensemble averaging sometimes improving output quality and efficiency at the same time.

Read the original